Todd Fisher traces a personal guitar project from a Halloween light display to a plugin that makes an instrument trigger spoken words. The project combines a digital audio workstation plugin with text to speech and note driven playback.
Todd Fisher explains why automatically slicing speech into words is difficult and compares silence based segmentation with syllable detection. He then uses pitch detection and a vocoder to map guitar notes onto speech, balancing the synthesized pitch with the voice signal.
Todd Fisher extends the prototype into a conversational loop: microphone input becomes text through speech recognition, a local language model generates a response, and speech is routed back through the guitar. Preprocessed vocal samples and pitch shifting move the project closer to singing, although the live demos expose latency, segmentation, and clarity limitations.
Watch on YouTube



