Why Does My AI Get Worse the More Context I Give It?

Share This Post
Why Does My AI Get Worse the More Context I Give It?

If you have been adding detail to your prompts, tools to your stack, and options to your offer and things keep getting muddier, this article answers the question nobody selling you software will raise: what if the fix is removal?


I Spent a Year Adding Things

I did an inventory in the spring. Every tool, offer, channel, and process I had added to my business in the previous twelve months.

The list was long and, honestly, a little impressive. Then I tried to write the second list, which was everything I had removed in the same period, and I could not fill a line.

That is the part I have been sitting with. Not that I had added too much, though I had. That I had never once treated removal as a legitimate move. Adding felt like progress and it produced something I could point at. Removing felt like admitting a mistake, and I have a lot of practice at not doing that.

Here is the direct answer to the headline question. Your AI gets worse with more context because models degrade measurably as input grows, even when the useful information is still in there. Stanford’s research documented accuracy drops above 30 percent when key information sits in the middle of a long context rather than at the edges. Chroma’s 2025 testing across 18 frontier models found every one of them degrading as input length increased, even with evidence held fixed and favorably placed. More context is not neutral. It buries the signal, and you pay for the burial twice: once in tokens and once in accuracy.

But the thesis of this article is bigger than prompts, because I have come to believe the same mechanic runs through everything. The skill that pays in 2026 is subtraction. Less in the context window. Fewer models in the critical path. A smaller audience described more precisely. And nobody is going to sell it to you, because there is no revenue in telling someone to have less.


Key Takeaways

  • Stanford’s lost-in-the-middle research found accuracy drops above 30 percent when key information is buried in the middle of a long context rather than placed at the start or end.
  • Chroma’s 2025 testing across 18 frontier models found performance degrading as input length grew even when the relevant evidence stayed fixed and well-positioned.
  • The largest ongoing discussion in the biggest prompt engineering community is that context compression matters more than prompt engineering.
  • Some of the highest-performing automations reported in production right now contain no AI model in the critical path at all.
  • Every credible AI educator has narrowed to one specific audience rather than teaching AI generally, which is the same discipline applied to positioning.

The Problem: Adding Is the Only Move Anyone Taught Us

Look at the incentives and the behavior stops being mysterious.

Every vendor in your stack is paid to sell you more seats, more features, more tiers. Every course promises additional capability. Every newsletter, including the good ones, is structurally a list of new things. I write content in this space and I am part of this. The format rewards addition.

So when something is not working, the trained response is to reach for another input. Prompt not producing what you want? Add more detail. Workflow unreliable? Add a step that checks it. Offer not converting? Add a bonus. Team not adopting the tool? Add a training. Each of those is a reasonable local move and together they produce a business that nobody, including you, can hold in their head.

I want to be specific about my own version, because the abstraction is too easy to nod along with.

My most-used prompt had grown to something like nine hundred words. It had accumulated over a year, one defensive clause at a time. Every time the output missed, I added an instruction to prevent that miss. By spring it contained two instructions that directly contradicted each other and had presumably been quietly fighting for months. I never noticed, because I never read the whole thing at once. I only ever appended.

Same pattern in the business. Tools nobody owned. Offers that overlapped enough to create decision friction for buyers but not enough for me to notice I had built a maze. Context files with January facts in an August world.

And here is the honest part about why I did not fix it sooner. Removing something requires admitting that adding it was a mistake, or at least that its usefulness expired. That is a small ego cost each time, and small ego costs are extremely effective at preventing action. It is much more comfortable to add a tenth thing than to look at the previous nine.

But what if the accumulation was the actual problem, and not the thing the accumulation was trying to fix?


The Evidence: The Research Is Unusually Clear

Four findings, and they point the same direction with more agreement than research usually offers.

Position beats volume, dramatically. Stanford’s lost-in-the-middle work found that models perform best when relevant information sits at the very start or very end of the context window, with accuracy following a U-shaped curve and dropping by 20 to 30 points or more when the same information is buried in the middle. Read that carefully, because it is counterintuitive. The information was present. The model had it. It still performed worse. Adding context does not just fail to help, it can actively hide what you already gave it.

Degradation happens at every length increment, in every model tested. Chroma’s 2025 research tested 18 frontier models including GPT-4.1, Claude Opus 4, and Gemini 2.5 and found all of them exhibiting this behavior at every input length tested. One controlled study cited in that body of work found reasoning accuracy falling from 0.92 to 0.68 as inputs grew from a few hundred tokens to three thousand. Not thirty thousand. Three thousand. That is a normal prompt with a document attached.

The practitioner community figured this out before the vendors admitted it. The largest ongoing discussion in the biggest prompt engineering community is titled, roughly, that context compression is probably more important than prompt engineering. Right behind it sits a post explaining that the cheapest optimization of the year was deleting website navigation from context entirely. These are not theorists. These are people watching their own bills and their own output quality.

The best automations in production have no model in them. This is the finding that generalized the principle for me. In the automation communities, the workflows people report actually running reliably in production include a missed-call text-back system for a mobile truck repair shop and a WhatsApp logistics dispatch workflow explicitly labeled as having no LLM node. Meanwhile the same communities are full of posts about agents that never declare themselves finished and automations that fail silently. The reliable ones are the simple ones.

And then the pattern I did not expect to find. I spent a morning looking at six well-regarded AI educators, people who genuinely know this material. Not one of them teaches AI. One teaches AI to course creators. One to women entrepreneurs. One to creatives doing brand imagery. One to Australian service businesses. One to introverted creators. One to visual marketers. Every single one narrowed, and narrowing cost each of them a large chunk of an available audience.

Same discipline. Different object.


What Changed for Me: Cutting Until It Broke

I did something with that nine-hundred-word prompt that I should have done a year earlier. I read the whole thing, out loud, in one sitting.

Then I found the one instruction buried inside it that actually described what I wanted, stated it in a single sentence, and rebuilt outward from there. I added back only the pieces that demonstrably changed the output, and I tested each one instead of assuming.

The rebuilt version is around three hundred and fifty words. The output is better. Not slightly better in a way I am talking myself into. Noticeably better, requiring less correction, and it costs roughly a third as much to run.

The method that made it work was cutting until quality actually dropped rather than cutting until I felt nervous. Those are wildly different thresholds and my nervousness fired somewhere around a 20 percent reduction. Real degradation did not appear until I had cut about 60 percent. That gap between where I got uncomfortable and where it genuinely broke was the entire opportunity, and I would have never found it without testing.

Two specific replacements did most of the work. First, I swapped description for example. I had three paragraphs describing the tone I wanted. One annotated example of my actual writing, with a short note about what to notice in it, outperformed all three paragraphs and used a fraction of the space. Second, I moved stable context out of the prompt and into a reference file, so the prompt stayed about the task rather than restating who I am every time.

Then I applied the same discipline outward, which was harder because there was no test I could run in an afternoon.

I removed a tool nobody owned. Nothing broke. I collapsed overlapping offers into a smaller set, and the thing I feared, losing the revenue attached to the removed options, did not happen, because buyers were not choosing between them so much as stalling in front of them. I pruned context files, cutting January facts and the two contradicting instructions.

The uncomfortable lesson underneath all of it: in every single case, I had been protecting something that was not producing anything. And I had been protecting it not because I had evaluated it, but because removing it would have required me to look at it.


Practical Steps

1. Read your longest prompt out loud in one sitting. Not skim. Read it. If it has grown over months by appending, you will find contradictions you have never noticed, because appending never requires you to look at the whole. This takes ten minutes and it is the highest-return ten minutes available.

2. Find the one sentence that is the actual instruction. Underneath every bloated prompt is a single request. State it in one sentence. Everything else in the prompt is context, hedging, or scar tissue from a past failure, and each piece has to earn its place separately.

3. Rebuild outward and test each addition. Add back only what demonstrably changes the output. Run the task with and without each piece. Most of what you add back will turn out to do nothing, which is information you cannot get by reasoning about it.

4. Cut until quality drops, not until you feel nervous. Those thresholds are far apart. Mine were 20 percent and 60 percent. The distance between them is where your savings and your accuracy gains both live, and you only find it by going past comfortable.

5. Replace description with one annotated example. Three paragraphs describing your tone will lose to one sample of your actual work with a note about what to notice in it. This is the single highest-leverage substitution in prompting and it also shortens everything.

6. Ask whether the model needs to be in the workflow at all. Go step by step through your most important automation and ask what the model is genuinely contributing. If the answer is that it sounds nicer, a template will do the same job and will not have a bad Tuesday.

7. Schedule subtraction, because addition happens on its own. Put a recurring review on the calendar for tools, offers, and context files. Ask one question about each: what has this produced since the last review. Addition never needs a reminder. Removal will never happen without one.


Frequently Asked Questions

Does a bigger context window mean I can stop worrying about context length?

No. Chroma’s 2025 testing found all 18 frontier models degrading as input grew, even when relevant evidence stayed fixed and well-placed. A million-token window means the model will accept a million tokens, not that it will reason well across them. Capacity and performance are separate questions.

Where should I put the most important information in a long prompt?

At the very beginning or the very end. Stanford’s lost-in-the-middle research found accuracy following a U-shaped curve, with performance dropping 20 to 30 points or more when key information sits in the middle. If something must be in a long context, place it at an edge.

How do I know if I have cut too much?

Run the same real task before and after and compare outputs against your actual purpose, not against your preference. Most people stop cutting at the point of discomfort, which arrives long before genuine degradation. Keep going until quality measurably drops, then add back one step.

Should I remove AI from workflows that are working fine?

Not because of ideology. Ask what the model contributes at each step. If it handles genuine variation that rules cannot, keep it. If it is producing nicer-sounding output that a template would match, a template is more reliable, cheaper, and unaffected by vendor outages.

Is narrowing my audience really the same principle as trimming a prompt?

Structurally, yes. Both are cases where value comes from exclusion rather than inclusion. A prompt that specifies everything gives the model nothing to prioritize. Positioning that includes everyone gives a prospect no signal that you are for them. Precision requires leaving things out.


Nobody Gets Excited About Buying Less

Two lists. One long, one empty.

I still think about that, because the empty list was not an oversight. It was the predictable output of a system where every incentive I encountered, including the ones I create for other people, pointed toward more. There is no vendor whose business model involves telling you to remove their product. There is no course called Have Fewer Things. I have never seen a conference talk on subtraction and I would probably not attend one.

Which is exactly why it works.

Everything in this business is competitive except this. Everyone has access to the same models, the same tools, the same pricing that keeps falling. What almost nobody will do is the unglamorous, mildly humiliating work of going back through what they built and taking things out. It requires admitting that some of what you added was never worth adding, and that particular admission is expensive in a currency that has nothing to do with money.

So do the thing nobody will sell you.

Read your longest prompt out loud tonight. Find the sentence that is actually the request. Cut everything until something breaks, and notice how much further past comfortable that turned out to be.

Then go find one thing in your business that nobody owns, and remove it, and watch nothing happen.

The work is not adding intelligence. Everyone has intelligence now. The work is clearing enough space that the intelligence can find what matters.


Jonathan Mast is the founder and CEO of White Beard Strategies. He works with entrepreneurs on using AI practically rather than performatively, serves a community of more than 100,000 business owners, and created the Perfect Prompt Framework. His most-used prompt is now three hundred and fifty words, down from nine hundred, and he is still slightly annoyed about how much better it works.