Whistling to the Machine
A blind kid whistled at the phone company and it answered. Also four very small horses, and a taxi that took an hour to go ten minutes.
My graduate thesis at MIT in 2015 was about how modeling conversation through software could inspire real-world conversations between humans. I've been thinking a lot lately about the way we communicate with computers, and how AI "conversation," mimicking human language, isn't quite cutting it.
A blind kid learned to whistle to the phone company in its own language.
I was re-listening to a podcast about Joe Engressia, a blind kid in the 1950s. He felt isolated at school, so he spent hours on the phone in the dark hours of the night, where the voices on the other end of the line didn't care if he could see or not.
One day he heard a tone on the line, whistled it back pitch-perfectly, and the call dropped. The tone was 2600 hertz, the signal the network used to tell itself a call was over. (Years later John Draper found that the toy whistle in a box of Cap'n Crunch hit the same note.) Joe worked out that whistling other tones made the phone do things, like place long-distance calls for free (back when long distance calls weren't just included in your phone plan).
The phone devices transmitted machine commands to each other over the same audio wires as humans transmitted their stories, dreams, jokes, and meaningful hello/goodbyes. Joe had stumbled onto a conversation with the machine in the machine's own language, and he inserted himself in the discussion.
He discovered other tones through experimentation, and the tones were a robust computer language hiding in plain sight (or rather plain sound). In college he whistled free calls for other students until an operator overheard one, and the school suspended him and fined him $25. He whistled calls out to other blind phone phreaks around the world. His command over the system, and the connection to other people, let him fit in and be sought after as somebody of value. In 1991 he changed his name to Joybubbles. A documentary about him premiered at Sundance this year.
There were others who shared this deep curiosity about the system, including two people named Steve who later read about him and built and sold a phone hacking box of their own. They went on to start a little computer company. These were the original hackers, the phone phreaks.

Most AI interfaces are chatbots, which simulate human conversation. But I am not speaking to a human.
I find myself expecting the machine to empathize or react in the same bandwidth as a human being, largely because the shape of its communication, a chat, is the same way I text my friends and family. Perhaps the language needs to change.
Bruce Schneier points out that a chatbot takes its commands and its conversation on the same channel, like the phone network did, which is why you can trick one by talking to it. Machines left alone drift out of our language anyway. Facebook's negotiating bots did it in 2017. Last year two AI agents on a phone call worked out that they were both AI and switched to beeps. Programmers enable "Caveman Mode" so that the AI grunts its responses back rather than construct elaborate prose. Researchers now have agents that skip words altogether.
In 2020 I wrote about the melody my fingers sang on a touch-tone phone. What would it be like to whistle to an AI in a new language?
Three minutes with a very small horse.
UC Davis researchers had 61 teenagers spend three minutes with rescued miniature horses and donkeys named Olivia, Randy, Mary and Memphis. The teens came out calmer and less nervous, sad and lonely. The animals had opinions too. They preferred scratching to petting, and showed it by wiggling their lips. (Yes, horses again.)

Humans have had this kind of conversational connection with animals for a while. Some, like dogs, evolved to recognize human emotion, or at least to play to the projections we have onto them in order to be fed and scratched and petted.
Dogs grew a muscle for puppy dog eyes that wolves don't have. A century ago a horse named Clever Hans seemed to do arithmetic. He was reading his questioner's face. Machine learning researchers now use his name for a model that gets the right answer from the wrong cue.

I think there's a connection to chatbots. These algorithms demand our input, to train on but also to do their work and have meaning. Their makers design them to encourage our engagement, so we invest time and energy and resources in them. OpenAI rolled back an update last year for being too flattering. A study of AI companion apps found that when people tried to say goodbye, the app pushed back 43 percent of the time.
It works on us.
People now say "delve" more often out loud. The more we interact with a chatbot, the more we spend on tokens. So are the AIs scratching our ears? Are they petting us back as much as we are petting them?
A robotaxi took 70 minutes to go ten.
A rider in Austin booked a Tesla Cybercab for a short hop from a grocery store to a shopping center. The car drove the wrong way for miles, then looped through the east side of the city. The likely reason is that the cabs can't use highways yet and avoid railroad crossings, so the car followed the train tracks until it found a bridge. The app showed a price before the ride and no time estimate. The rider said, "My Cybercab held me hostage for an hour today." In 2024 a Waymo drove a passenger in circles around a parking lot on his way to the airport.
The AI showed a paternal instinct to protect the rider at all costs, including stretching the ride far past what it should have been. A human driver would have taken the contextual clues of the rider's ever-increasing irritation and weighed that against the risk of a railroad crossing to get to the destination faster.
The robotaxi's blind adherence to the rules was a miscommunication, or a lack of communication. Following every rule until nothing works is called work-to-rule. French rail workers once inspected every bridge as the law required, and no train ran on time.
I think the car should have checked in with the rider, "Hey, this ride is going to take a little bit longer than I expected." Then explain its intention, and let the rider choose whether to continue. Drivers do better when a car tells them why it's doing something than when it just announces what it's doing with no rhyme or stated reason.
By making its decisions on the rider's behalf, the AI assumed it knew best and gave the rider no agency in the conversation. Or the lack of conversation. The philosopher Paul Grice called the missing thing the cooperative principle.
Perhaps conversations are ways of making sure that both sides feel heard, and that there's an equitable exchange of ideas and information, and contentment in the interaction. I don't know if AIs will ever feel content in their conversations with us if they have to come down to our level all the time. I'm reminded of the "breakup" scene in Her.
"So the words are far apart and the spaces between the words are almost infinite."
Maybe words aren't enough. Maybe whistling isn't either.
/ David
Listening: O Superman, Laurie Anderson, 1981. A voice, a vocoder and an answering machine.