Why Does It Take Me Three Hours to Write Something I Could Explain in Ten Minutes?

Share This Post
Why Does It Take Me Three Hours to Write Something I Could Explain in Ten Minutes?

Because your fingers are the bottleneck, not your brain. Here is what the research actually says about the gap between talking and typing, and why the tools that close it finally got good this month.


I spent about thirty years assuming I was a bad writer.

Not a bad thinker. I could explain a strategy on a phone call in eight minutes and the other person would say "that was really clear, can you send that to me?"

Then I would open a blank document and sit there for two hours producing four paragraphs I did not like.

I blamed discipline. I blamed my ADHD. I blamed the fact that I taught myself to type and still hunt for the number keys.

Here is the thing. It was none of those. The answer is boring and mechanical.

You think and speak at roughly 150 words per minute. The average person types at 52. Your first draft is not slow because your thinking is slow. It is slow because you are forcing 150 words per minute of thought through a 52 word per minute pipe.

That gap has existed since the invention of the keyboard. Dictation was always the theoretical fix, and it was always useless in practice, because cleaning up the transcript took longer than typing would have. Anyone who tried Dragon in 2009 knows exactly what I mean.

That tax is gone now. Whisper-class speech models killed it, and this week two independent things landed 24 hours apart that make the case impossible to ignore.

Michael Hyatt published a piece on August 25 in his AI Business Lab newsletter calling the typing gap a tax on your own thinking, and naming the dictation tools that converted him. The next morning, Google DeepMind shipped Gemini 3.5 Transcribe, a model built specifically to turn rambling speech into clean formatted text.

Neither was referencing the other. That is usually the signal that something real happened.

So here is my thesis, plainly. If your first draft has been the bottleneck rather than your thinking, you can remove that bottleneck this week for less than the cost of lunch. And most of you have been misdiagnosing the problem as a character flaw.


Key Takeaways

  • Research puts average conversational speech near 150 words per minute and average typing at 52, which makes your keyboard roughly a 3x drag on your own thinking.
  • Dictation failed for decades because of the editing tax, not the speed math, and modern speech models have largely eliminated that tax.
  • The bottleneck for most business owners is the first draft, not the idea, and those are two completely different problems with two completely different fixes.
  • A working voice-first system has three layers: capture your speech, let AI clean the transcript, then apply your own judgment to what comes back.
  • If you have ADHD, dyslexia, or you simply never learned to type, this is not a productivity hack. It is removing a tax you have been paying your whole career.

The Problem Nobody Names Correctly

Ask a business owner why they do not publish more and you will hear the same three answers. No time. Not a good writer. Do not know what to say.

The third one is occasionally true. The first two are almost always a misdiagnosis.

Watch what actually happens. You know the point you want to make. You could make it out loud, right now, to a friend at a coffee shop, in about six minutes, with better examples than you will ever put on paper.

Then you sit down to write it and something strange happens. You lose the thread. You rewrite the opening line four times. Forty minutes later you have a paragraph and a vague sense that you are not smart.

That is not a writing problem. That is a bandwidth problem.

Ronald Kellogg, a cognitive psychologist who spent his career studying how people actually compose text, describes writing as three processes running at once: planning what you want to say, translating that into actual sentences, and reviewing what you already produced. All three pull from the same limited pool of executive attention.

Translating is the expensive one. And here is the cruel part. While you are spending your working memory on spelling, punctuation, and hunting for the semicolon, the idea you were holding starts to leak out.

You have experienced this. You knew exactly what you meant, and by the time you finished the sentence you had forgotten the better version of it.

I am wired for hyper-focus. When I lock in on something, I can go five hours without looking up. So for years I could not square why writing, of all things, was the one activity where I could not stay in the tunnel.

Candidly, it took me an embarrassingly long time to figure out that the tunnel was fine. The doorway was too narrow.

I have been broke, I have been divorced, and I have been to prison. None of those taught me to type any faster. I have made peace with three fingers.

What I had not made peace with was the assumption that the three fingers were the ceiling on what I could produce.


What the Research Actually Says

I want to be careful here, because this topic attracts made-up numbers. Here is what is actually documented.

1. Speech is about three times faster than typing, and more accurate.

In 2016, researchers from Stanford, the University of Washington, and Baidu ran a controlled experiment putting speech recognition against 32 people aged 19 to 32 typing on an iPhone. These were not slow typists. As Stanford professor James Landay put it, "they grew up texting, so we're putting speech recognition up against people who are really good at this task."

Speech was 3.0 times faster than the keyboard in English, with an error rate 20.4 percent lower. In Mandarin it was 2.8 times faster with an error rate 63.4 percent lower.

That study published on August 25, 2016. Michael Hyatt's newsletter published on August 25, 2026. Ten years to the day, and the finding is only now becoming usable by normal people.

2. The average person types 52 words per minute.

Researchers at Aalto University and the University of Cambridge analyzed 136 million keystrokes from 168,000 volunteers, the largest study of everyday typing ever conducted, presented at the CHI conference in 2018.

Average speed was 52 words per minute. Most people landed between 30 and 60. The fastest hit 120.

3. Conversational speech runs about 150 words per minute.

The National Center for Voice and Speech puts the average conversation rate for American English speakers at roughly 150 words per minute. Most adults sit somewhere between 120 and 150 in ordinary conversation.

Do the math yourself. 150 divided by 52 is about 2.9.

Hyatt framed it as a 4x tax. The verified numbers put it closer to 3x. I am going to give him the framing and correct the multiplier, because a real 3x is more persuasive than an inflated 4x, and because you should be suspicious of anyone quoting numbers they cannot source.

4. Transcription accuracy stopped being the problem.

In OpenAI's 2023 Whisper paper, the best zero-shot model reached a 2.5 percent word error rate on the LibriSpeech clean test set, roughly matching human transcribers on the paper's own comparison. Not the best human. Comparable to a human.

That is the number that changed everything. When the machine is as accurate as a person, the editing tax disappears.

5. The new models clean up your speech, not just capture it.

Gemini 3.5 Transcribe shipped on August 26, 2026. Per Google's announcement and Engadget's coverage, it automatically detects more than 85 languages, removes filler words, resolves self-corrections, applies structured formatting, and attributes speech to up to three speakers with word-level timestamps.

Read that middle part again. It resolves self-corrections. When you say "we should hire two people, actually three," it knows you meant three.

That is the piece that was missing. Old dictation transcribed sound. These models transcribe intent.


The System: Capture, Clean, Judge

I think like a systems architect, so let me lay this out as three layers instead of a list of apps. Get the layers right and the specific tools become swappable.

Layer one: capture.

The only job here is to get your spoken words into text with as little friction as possible. Friction kills this. If it takes four clicks to start dictating, you will not do it.

Michael Hyatt names Wispr Flow as the tool that converted him and Willow Voice as his current pick. Both run system-wide, meaning you hold a key, talk, and the text appears in whatever field your cursor is in. Email, Slack, Google Docs, your CRM. Both run around $12 to $15 a month.

Google says Gemini 3.5 Transcribe is coming to any text field in Chrome, plus Docs, Keep, and Gmail. If you live in Google's world, that may end up being free and good enough.

And the dictation already built into your phone and your Mac is dramatically better than it was two years ago. Start there before you spend a dollar.

Layer two: clean.

This is where the editing tax used to live and where AI now does the work. You take the raw transcript and hand it to Claude or ChatGPT with instructions to strip filler, resolve your self-corrections, and add paragraph breaks.

Thirty seconds. Sometimes less.

The critical instruction, and I mean critical, is telling it to keep your words. Left alone, these models will smooth your voice into corporate oatmeal. You want the transcript cleaned, not rewritten.

Layer three: judge.

This layer is you and it is not optional.

The machine cannot tell you whether the argument holds, whether the example lands, or whether you should have said the harder true thing instead of the easier comfortable one. That is judgment, and judgment is the part of your work that is actually worth something.

Here is what most people get wrong about AI. They try to hand it layer three. Then they are disappointed when the output is generic, and they conclude AI does not work for them.

AI is not there to replace your thinking. It is there to remove the tax you pay to get your thinking out of your head. That is the whole game.


How to Test This in the Next 48 Hours

Do not overhaul your workflow. Run one experiment.

  1. Pick something you already owe someone. A follow-up email you have been dreading. A proposal. A post you have started three times. Choose something with real stakes, not a practice piece, because you need to feel whether the output is usable.

  2. Turn on the dictation you already have. On a Mac, press the dictation key. On a phone, tap the microphone. Do not buy anything yet. The goal today is to test the concept, not the product.

  3. Talk it out for five minutes without stopping. No backing up, no restarting sentences, no editing while you speak. If you correct yourself, just keep going and say the right version. The cleanup layer will handle it. This will feel deeply wrong for about ninety seconds.

  4. Time yourself, then look at the word count. Most people produce 600 to 800 words in five minutes. Sit with that number and compare it honestly to how long the typed version would have taken.

  5. Hand the transcript to your AI with this prompt. Fill in the brackets with your own details:

[The Job]
Clean up this raw dictation into a first draft I can edit.

[The Background]
This is a spoken transcript about [WHAT YOU WERE TALKING ABOUT]. I was thinking out loud, so it contains filler words, false starts, and places where I corrected myself mid sentence. The audience is [WHO YOU ARE WRITING THIS FOR]. My natural tone is [DESCRIBE HOW YOU ACTUALLY TALK, FOR EXAMPLE: plain, direct, a little dry]. Here is the transcript: [PASTE TRANSCRIPT]

[The Deliverable]
Return the cleaned draft with filler words removed, self corrections resolved to the version I landed on, and paragraph breaks added. Keep my word choices and my sentence rhythm. Do not add ideas I did not say and do not make it sound more formal than I am. At the end, list anything that was unclear.

[The Questions]
Ask me any questions you have.

  1. Edit the result and send it. Actually send it. The experiment does not count if you keep it in a drafts folder. You are testing whether this produces work you will stand behind.

  2. Repeat it four more times before you judge it. The first attempt is always awkward because you are fighting the habit of composing in your head. By the fifth, it starts to feel like the natural way to work.


Frequently Asked Questions

Does dictation actually work if I have an accent?

Yes, far better than it used to. Whisper-class models were trained on enormous amounts of varied real-world audio rather than clean studio recordings, which is exactly why they handle accents, background noise, and imperfect microphones so much better than the dictation software you gave up on years ago.

Will my writing sound like AI wrote it?

Only if you let the AI rewrite instead of clean. The words come from your mouth, so the ideas, examples, and phrasing are yours. Explicitly instruct the model to preserve your word choices and sentence rhythm, then read the output and put back anything it smoothed away.

What if I am not comfortable talking out loud to a computer?

Nearly everyone feels ridiculous for the first few sessions. It helps to imagine you are explaining the thing to one specific person you like. Most people cross over somewhere around the fourth or fifth attempt, at which point going back to typing feels like a downgrade.

Which tool should I actually buy?

Test your device's built-in dictation first, since it costs nothing and may be sufficient. If you want a system-wide tool, Wispr Flow and Willow Voice both run roughly $12 to $15 a month. Google is rolling Gemini 3.5 Transcribe into Chrome, Docs, and Gmail, which may cover you free.

Is this only useful for writing long content?

No, and honestly the biggest wins are small. Emails, Slack replies, meeting notes, CRM entries, text for your bookkeeper. The tasks you postpone because typing them feels tedious are exactly the tasks that vanish when talking becomes an option.


You Are Not a Slow Writer

I want to go back to where this started, because I think a lot of you are carrying something you do not need to carry.

You have a story about yourself. It probably sounds like "I am just not a writer." Maybe a teacher planted it. Maybe you compared your rough draft to somebody's finished book and drew the obvious wrong conclusion.

I carried that story for three decades. I built a business, spoke on stages, and taught hundreds of thousands of people, all while quietly believing that the writing part of my work was a weakness I was working around.

It was never a weakness. It was a transmission problem. I had the engine. I was hauling gravel in a wheelbarrow with a dump truck sitting right there in the driveway.

That is what makes this week worth paying attention to. Not because a new model shipped. Models ship every Tuesday. It matters because the specific thing that made dictation useless, the editing tax, is genuinely gone, and a lot of people are still avoiding the tool because of how badly it burned them in 2011.

So here is my invitation. Do not take my word for it and do not take Michael Hyatt's. Take the next thing you were going to write and talk it out instead.

Five minutes. One thing you already owe somebody.

If it works, you did not discover a productivity hack. You discovered that the ceiling you have been living under was never yours.

You were always able to say it. Now you can finally send it.

P.S. If you have ADHD, dyslexia, or you just never learned to type properly, you have been paying this tax with interest your entire career. Nobody handed you a discount. Go take it back.


About the Author

Jonathan Mast is the founder of White Beard Strategies, where he teaches non-technical business owners to use AI to amplify skills they already have rather than replace them. He leads a Facebook community of more than 500,000 members and the AI Insiders membership, and he speaks regularly on practical AI adoption for small business.

He types with three fingers, has never once considered fixing that, and now dictates most of his first drafts.


Sources