Notes on keeping a journal · · 7 min read
Voice journaling: how to talk your entries out without hating the transcript
Why speaking into a phone is about three times faster than typing, where the dislike of your own voice comes from, what research says about talking it out, and how on-device transcription works.
- voice journal
- voice notes
- voice transcription
- journal
- privacy
In 2018 a group from Stanford and the University of Washington sat people down with a phone and asked them to enter the same short messages two ways: with their fingers on the on-screen keyboard, and by voice through speech recognition. In English and in Mandarin, measuring speed and the errors left after correction.
In English, voice came out roughly three times faster than the keyboard; in Mandarin, 2.8 times. And the finished text had fewer errors with voice, not more. This was measured on short phrases, not journal entries, but the ratio carries over: in one minute you say what you would have typed in three.
And still most people I know do not want to talk their entries out. Not because of speed. Because afterwards you have to listen to it, or read it.
Why your own voice is unpleasant
That feeling has a name and a date. In 1966 Philip Holzman and Clyde Rousey described "voice confrontation": people were played recordings, one of which was their own, and many reacted to their own with puzzlement or dislike, sometimes without recognising it at all. The explanation has two layers. The first is physics: from the inside you hear yourself partly through the bones of the skull, which add low frequencies, so on a recording the voice sounds higher and thinner than it "should". The second layer is more interesting: a recording lets you hear what you normally do not notice while speaking. Hesitation, haste, an intonation you do not agree with.
A transcript does the same thing on paper. Spoken language looks bad as text. "Well", "sort of", repetitions, sentences that start and never finish. Typed text you edit as you go, and it comes out smoother than the thought. A transcript shows the thought as it was, with a stumble exactly where you did not actually know what you thought.
That is the trap. You want to fix or delete the transcript because it is "ugly". But the ugliness is the content of the entry. The pause before the word "fine" in answer to "how was your day" carries more information than the word.
Saying it out loud works no worse than writing it
Most of what is known about the benefits of journaling comes from writing: a person writes about something hard for twenty minutes on several consecutive days. But two studies also tested voice.
In 1994 Brian Esterling and colleagues split students into three groups: some wrote about a stressful event, some told it out loud into a recorder, and some wrote on neutral topics. What they measured was not mood but a biological marker of strain on the immune system. Changes appeared in both disclosure groups, and in the speaking group they were larger than in the writing group. A small sample, students, one study, so it does not follow that voice is "better". What follows is that the medium matters less than the act of turning an event into words.
In 2006 Sonja Lyubomirsky and colleagues compared three ways of dealing with the worst event of one's life: writing about it, talking about it, and simply thinking about it. Writing and talking helped; thinking did not, and sometimes made things worse. I go into that study more in the piece on why we quit journaling; the point here is a single one: talking into a phone and writing in a notebook land in the same category, and silently replaying it in your head lands in another.
Writing or talking about a hard event helped; thinking about it without putting it into words did not.
How to talk it out so you do not erase it later
Imagine someone, call him Kostya, who has not written a line in a journal in three years because "there is nothing to write". Six months ago he started hitting record on the walk from the subway to his door, nine minutes. He talks to no one in particular, straight ahead, as if thinking aloud. He does not reread the transcripts for weeks. Then he searches for the word "tired" and sees it in eighteen entries out of thirty, always on a Tuesday. The story is made up, but the mechanism in it is real: a voice entry pays off at the moment it can be found.
What follows from that mechanism in practice.
Any first sentence will do. Do not rehearse an opening. "So" or "today" is enough. The first ten seconds are a run-up anyway, and nobody grades them.
Talk ahead, not to a listener. The moment an imaginary audience appears, so does shame. It helps to picture yourself thinking out loud in an empty room. That is what a journal is.
One minute, maybe two. Longer and you start repeating yourself. If there is more to say, a second entry beats a ten-minute first one.
Do not listen back straight away. Right after recording your voice is still in your ears, and Holzman and Rousey's effect kicks in. A week later you are no longer listening to a voice but to the person who said it.
The transcript is a draft for search, not a text. Its job is for a word to be findable a month from now. It does not need to be pretty.
Where voice beats text
In the morning with your eyes closed, before the dream falls apart: there is a separate piece on dream journaling about that. On the road, when your hands are busy. When you are angry: words come faster than fingers, and the entry comes out more honest than if you gave yourself time to phrase it. When "there is nothing to write": talking about an empty day is easier than writing about it, because you already know how to talk, and writing takes gathering yourself.
Where voice is worse: in a room with other people, in a queue, and when you want to find the exact word. That is what text is for.
Where the voice goes
Now the part usually printed in small type. Speech recognition on a phone comes in two kinds: on the device itself and on a server. Apple's developer documentation has a switch that requires processing on the device only; with it on, audio is sent nowhere, and if local recognition is unavailable on that phone the request simply fails instead of quietly going to the network.
In Lanternly that switch is always on. Voice is recorded and transcribed on the iPhone, and the transcript has no network path. Three honest consequences follow.
First: local transcription runs through the system dictation of iOS. If dictation is switched off in the phone's settings there will be no text, and the app says so plainly rather than showing an endless "transcribing". It is turned on in Settings, under General, then Keyboard. Second: the recognition language follows the app's interface language, English or Russian. Third: the local model is worse than a server one on names, jargon and whispering. The audio stays in the entry regardless, and transcription can be run again.
If where your voice recordings go matters to you, the private journal app checklist has an item on exactly that: ask the app where the voice is transcribed, not only where the text is stored.
One minute
Record one minute today. Do not listen to it. A week from now open the transcript, find the clumsiest spot in it and look at what it is about. Most likely that is exactly where the thing worth recording was.
Sources
- Ruan, S., Wobbrock, J. O., Liou, K., Ng, A., Landay, J. A. Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 1(4), 1–23, ACM (2018)
- Holzman, P. S., Rousey, C. The voice as a percept. Journal of Personality and Social Psychology, 4(1), 79–86, American Psychological Association (1966)
- Esterling, B. A., Antoni, M. H., Fletcher, M. A., Margulies, S., Schneiderman, N. Emotional disclosure through writing or speaking modulates latent Epstein-Barr virus antibody titers. Journal of Consulting and Clinical Psychology, 62(1), 130–140, American Psychological Association (1994)
- Lyubomirsky, S., Sousa, L., Dickerhoof, R. The costs and benefits of writing, talking, and thinking about life's triumphs and defeats. Journal of Personality and Social Psychology, 90(4), 692–708, American Psychological Association (2006)
- requiresOnDeviceRecognition. Speech framework documentation, Apple Developer