Bigger Context Windows Didn't Make Our RAG Smarter —…
    Neura MarketNeura Market/Perplexity
    ChatGPTChatGPTClaudeClaudeGeminiGeminiCursorCursorGrokGrokPerplexityPerplexityDeepSeekDeepSeek
    CoPilotCoPilotStable DiffusionStable DiffusionMidjourneyMidjourney
    View All Directories
    OverviewRulesPromptsMCPsAgentsGamesBlogVideosGuidesCoursesCommunityTrending
    PerplexityBlogBigger Context Windows Didn't Make Our RAG Smarter
    Back to Blog
    Bigger Context Windows Didn't Make Our RAG Smarter
    ai

    Bigger Context Windows Didn't Make Our RAG Smarter

    ValeryKot July 8, 2026
    0 views

    We stopped measuring retrieval quality by how many tokens we could fit into the prompt. When...

    We stopped measuring retrieval quality by how many tokens we could fit into the prompt.

    When long-context models became available, many of us made the same assumption.

    If an LLM can read 128K tokens, retrieval suddenly feels less important. Why spend time carefully selecting documents if the model can simply read everything?

    It sounds reasonable.

    In practice, it wasn't.

    More context, worse answers

    Imagine asking your internal assistant:

    Why did we abandon microservices?

    Retrieval returns thirty documents.

    • an architecture decision record
    • a few Jira tickets
    • Slack discussions
    • meeting notes
    • a glossary page

    Everything is related.

    Almost nothing answers the question.

    The actual decision lives in a single ADR written months earlier. It explains the trade-offs: team size, latency, deployment complexity, operational cost.

    But that document isn't especially similar to the query. It doesn't repeat the same vocabulary. It doesn't even mention "microservices" very often.

    So it gets buried.

    The model now receives thirty relevant documents and does what language models are very good at: it produces a coherent explanation.

    The problem is that coherence is not the same thing as faithfulness.

    Instead of recovering the original decision, it often synthesizes one from recurring themes across the retrieved documents.

    The answer sounds plausible.

    It just isn't the answer that was originally made.

    Bigger windows don't fix retrieval

    Research has already shown that models struggle with information buried inside very long contexts. The Lost in the Middle paper is probably the best-known example.

    Our experience suggested something slightly different.

    Sometimes the answer isn't lost because the context is long.

    It's lost because the retrieval stage couldn't distinguish the document that contains the decision from documents that merely discuss the same topic.

    Adding more context doesn't necessarily solve that problem.

    Sometimes it simply gives the model more material to average together.

    We were optimizing the wrong thing

    For a while we treated retrieval as a packing exercise.

    How many useful chunks can we fit into the prompt?

    Over time the question changed.

    Why is this document here?

    Should it be here at all?

    Does it explain the decision, or does it merely mention the same technology?

    Those questions turned out to matter much more than the size of the context window.

    Retrieval is a selection problem

    The biggest shift wasn't moving from 8K to 128K tokens.

    It was realizing that retrieval isn't about fitting more information into a prompt.

    It's about selecting the few pieces of information that actually explain the answer.

    Large context windows are incredibly useful.

    They just don't compensate for weak retrieval.

    If anything, they make weak retrieval look convincing.


    Next time I'll look at another assumption I no longer believe: that documents should be treated as bags of chunks.

    Tags

    aillmmachinelearningrag

    Comments

    More Blog

    View all
    Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught megemma

    Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me

    I ported the whole Gemma-4 family — E2B, E4B, 12B, 31B, and the 26B-A4B MoE — to run on...

    X
    xbill
    Hey DEV, I'm Tobore. Let's actually connect.community

    Hey DEV, I'm Tobore. Let's actually connect.

    Hey DEV, I'm Tobore. Let's actually connect. I've been on here for a while now, mostly writing and...

    L
    Laurina Ayarah
    I burned through thousands of AI tokens. Then a friend did it for freeai

    I burned through thousands of AI tokens. Then a friend did it for free

    (yep, kinda clickbait, just for the funsies 😊) At the beginning of the year, I relaunched my...

    P
    Paulo Henrique
    Claude might be saturating your machineai

    Claude might be saturating your machine

    My laptop was sitting idle with the fan at full tilt. Nothing was running that I knew of. The culprit...

    S
    Sidhant Panda
    Automated GitHub Code Reviews Using Google Geminigithubactions

    Automated GitHub Code Reviews Using Google Gemini

    I Built a Thing! TL;DR — Google Gemini-based Pull Request reviews and Issue Triaging for...

    D
    Darren "Dazbo" Lester
    What is an "agentic harness," actually?ai

    What is an "agentic harness," actually?

    I've been hearing the word "harness" thrown around a lot lately. I assumed it just meant "the IDE" or...

    T
    Tilde A. Thurium

    Stay up to date

    Get the latest Perplexity prompts, rules, and resources delivered to your inbox weekly.

    Neura Market LogoNeura Market

    Discover the best AI prompts, plugins, and resources for Perplexity and more.

    Content Types

    • Rules
    • Prompts
    • MCPs
    • Agents
    • Guides

    Platforms

    • ChatGPT Directory
    • Claude Directory
    • Gemini Directory
    • Cursor Directory
    • Grok Directory
    • Perplexity Directory
    • DeepSeek Directory
    • CoPilot Directory
    • Stable Diffusion Directory
    • Midjourney Directory
    • All Directories

    Resources

    • Blog
    • Documentation
    • Help Center
    • Marketplace

    Legal

    • Privacy Policy
    • Terms of Service

    © 2026 Neura Market. All rights reserved.

    |

    Not affiliated with any AI platform vendors.

    Neura Market

    Custom AI Systems & Services

    Our team of experienced AI builders will help build custom AI systems, workflows, and solutions for your business.

    Request custom work

    Ready-made automations for this

    Workflows from the Neura Market marketplace related to this Perplexity resource

    • Efficient Calendar Slot Retrieval for Project Managementmake · $19.24 · Uses make
    • Automate Content Creation with Google Sheets, Perplexity AI, and OpenAImake · $4.99 · Uses make
    • Automate Chat Responses from New Google Sheets Entries Using Perplexity AI and ChatGPTmake · $4.99 · Uses make
    • Automate Facebook Post Creation with Google Sheets and Perplexity AImake · $4.99 · Uses make
    Browse all workflows