Positioning note

Where context wins

Of all the ways a company can improve LLM output — evals, harness, context, post-training — which companies should put nearly everything into the context layer, and what that implies for who to sell to.

Internal discussion draft · August 2026

The sorting principle

The tempting split is facts versus skill: facts go in context, skill gets post-trained. That split is wrong, and wrong in the way that matters most — because a great deal of what an organization needs from a model is judgment, and judgment lands on both sides of the line.

"How we decide things here" is judgment, and it belongs in context. "Which of these deals smells like it will churn" is also judgment, and it belongs in weights. Sorting them requires taking judgment apart first.

Judgment is not one thing

ComponentExampleStatable?Home
Criteria — what mattersConcentration risk outweighs growth rateUsuallyContext
Thresholds — where the line sitsDiscounts over 20% need VP sign-offYes, and volatileContext
Weighting — relative importanceWeight retention twice logo countSometimesEither
Pattern recognitionThis one resembles the accounts that churnedRarelyWeights
Register — how output should feelHouse voice, hedging disciplineNo, but demonstrableWeights

Four of the five are statable. Only pattern recognition genuinely wants weights. That decomposition is why the thesis survives the obvious objection: once you take judgment apart, most of it turns out to be context-shaped.

Three tests, in priority order

  1. Can it be stated, or only demonstrated?

    If a competent owner can write the rule in a sentence, it belongs in context. If the honest answer is "show me five hundred examples and you'll develop a feel for it," you are fitting a function nobody can specify — that is a weights problem.

  2. What is its half-life?

    The volatility test applies to judgment exactly as it applies to facts. Risk appetite shifts quarterly. Bake it into weights and you have frozen last quarter's posture into an artifact you cannot inspect.

  3. Must someone be able to cite, challenge, or override it?

    If yes, it lives in context regardless of how the first two tests came out. Weights cannot be cited.

The third test is a veto, not a tiebreaker. Judgment that is stable and genuinely tacit still has to be externalized as a retrievable artifact if it must be defensible — you accept a worse mechanism to buy auditability. In regulated work this single test decides most cases before the other two are consulted.

The asymmetry that settles the rest

Context can emulate weights. Weights cannot emulate context. A rubric plus a handful of examples approximates post-trained judgment — imperfectly and expensively, but it works. Nothing approximates contextual judgment with weights, because what you need is instance-specific and permission-bound.

Context is therefore the general mechanism, and post-training is a compilation step: what you reach for when a piece of judgment is invoked on every call and costs too much to keep re-injecting. It is an optimization, not a knowledge strategy. Read that way, "spend 90% on context" stops being a claim about which lever is more powerful and becomes a claim about which one is load-bearing.

One desk, four pieces of judgment, three homes

"Restaurants in their first year are high risk"
Context — statable, and must be defensible to a regulator
"Our risk appetite this quarter"
Hot-injected context — applies to every call, but changes too fast to bake
"This applicant feels off in a way I can't name"
Weights — tacit, and they hold fifty thousand scored outcomes to learn from
"How a decline letter should read"
Split — required disclosures in context, tone in the system prompt

Same lending desk, four pieces of judgment, three different destinations. The question was never facts versus judgment.

The gate almost nobody clears

Post-training judgment needs a labeled corpus of your own decisions with scored outcomes — thousands of them, ground truth attached. Lenders have it. Insurers, ad platforms and high-volume support organizations have it. Almost nobody else does. Most companies weighing whether to fine-tune their judgment could not do it if they decided to.

The two ways this goes wrong

Baking what should have been written. Judgment goes into weights; later someone asks why the model said what it said and there is nowhere to look. The judgment becomes unlocatable and unfixable. In regulated work that isn't a performance problem, it's an audit finding.

Injecting what should have been baked. Forty pages of style guide in every prompt. Attention dilutes, adherence drifts, cost scales with every call. This is where post-training legitimately wins, and where context maximalists lose.

Context deserves the overwhelming share of effort when this decomposition keeps landing on the same side. In 2026 it usually does — frontier models are rarely too dumb for enterprise knowledge work. They are uninformed, and increasingly, they are uninformed about how the organization decides rather than merely what it knows.

The characteristics

  1. Truth has a short half-life

    The sharpest single test. If the correct answer to "what's our position on X" changes monthly, fine-tuning bakes in staleness faster than you can retrain. Ask: what's the half-life of a correct answer here? Weeks means the context layer is the only viable lever. Years means they should fine-tune and stop taking your calls.

  2. Truth is permission-differentiated

    Two employees ask the same question and correctly receive different answers based on what they're cleared to know. This rules out post-training outright — weights cannot be need-to-know — and it is the hardest property to bolt on after the fact. Companies with real internal compartmentalization aren't merely good prospects; they're companies for whom the alternatives are structurally unavailable.

  3. Knowledge is scattered and unwritten

    If everything already lives in one wiki, competent search wins and a context platform is overkill. The 90% case is knowledge spread across chat, email, documents, meetings, and a handful of people's heads — where the real work is reconciling it, not indexing it.

  4. Wide answer surface, low per-question difficulty

    A long tail of questions, each trivially answerable if you knew the fact. Contrast with novel reasoning work — drug discovery, hard code generation — where capability is the binding constraint and context is merely table stakes.

  5. Context is the durable asset, not just the bottleneck

    Both conditions must hold, not either. Some companies are context-bottlenecked but their context is one afternoon of digitization away from solved. That's low value and it churns. You want a bottleneck whose fix compounds.

  6. Organizational churn

    Fast headcount growth, turnover, reorgs, heavy contractor mix. These companies hemorrhage context structurally, and they already feel it as onboarding pain — which means the budget line exists before you arrive.

Why 90% rather than 40%

The asymmetry is depreciation. Model quality arrives free from the labs, so every dollar of post-training races a curve moving underneath you. Harness is commoditizing into frameworks. Both get cheaper to skip each quarter.

The context layer is the only lever that compounds, stays proprietary, and survives a model swap intact. Spending on the others is renting. Spending on context is owning.

One correction to the premise

Nobody should actually run 90/0 on context versus evals. You cannot steer context work blind — "we added more documents and it feels better" is how these programs die in month four.

The honest framing is 90% of improvement energy on context, with evals as the instrument rather than a competing investment. Practically: a prospect with no measurement discipline isn't disqualified, but you have to bring the ruler, and that's a real cost in the sales motion worth pricing in.

Disqualifiers

Discovery questions

  1. What's the half-life of a correct answer here — weeks, or years?
  2. Can two employees correctly get different answers to the same question?
  3. The last time someone told a customer something wrong: what did they not know, and where did that fact actually live?
  4. What happens to your answer quality when someone leaves?
  5. You've already tried a chatbot on your documents — what specifically disappointed you? If the answer is "it made things up about us," that's a context failure and they've pre-qualified themselves.

The profile in one line

Mid-size, knowledge-dense, permission-sensitive organizations — roughly 100 to 2,000 people, growing or churning — where the correct answer changes monthly, differs by who is asking, and currently lives in people rather than systems.

Regulated-adjacent sectors and professional services fit the shape well. So does any company whose product is accumulated judgment.

What the market data actually says

The frame above is argued. This section is sourced — and it moves two of the six characteristics from plausible to evidenced. Claims are tagged sourced where they trace to primary reporting or named research, and soft where the only available figures come from vendor marketing or unsourced aggregation.

The gap in the market

Glean began by selling to technology companies in the 500–2,000 employee range, then moved decisively upmarket. What it left behind is the relevant fact for anyone entering now.

softReported entry economics sit near a 100-seat minimum and roughly $60K annual contract floor, with all-in cost of ownership for mid-to-large deployments quoted between $350K and $480K per year. These figures come largely from competitor marketing content rather than Glean's own disclosures, so treat the magnitudes as indicative and the direction as reliable: the category leader has a floor, and it sits above most of the mid-market.

The result is a barbell. Large enterprises buy a horizontal context platform. Everyone below the floor assembles something out of wiki tools, bundled assistants, and in-house retrieval projects — which is precisely the cohort the six characteristics describe.

What they're using today

Any prospect in this band already has something. The question in discovery is never "do you have a knowledge tool" — it's which of these they've settled for, and which failure they've already felt.

What they runWhy they chose itWhere it breaks
Microsoft 365 Copilot Bundled, no procurement, already licensed Grounding against SharePoint and permissions is the most-cited reason its agents disappoint; users expect a consumer chatbot and get an assistant that doesn't know the company
Glean, Coveo, Sinequa Genuine cross-system breadth; the only category that doesn't force an ecosystem choice Seat minimums and total cost put them out of reach below roughly 1,000 employees
Confluence + Rovo AI bundled into every paid tier at a few dollars per user; already the engineering wiki Only as good as what someone wrote down, and stays largely inside Atlassian's own surfaces
Guru Named expert re-verifies each card on a schedule — the most conservative answer to "is this still true" Human curation doesn't scale to a wide answer surface; knowledge has to fit the card shape
Notion AI Zero friction if the company already lives in Notion Knows Notion. Does not know Slack, email, the CRM, or the meeting nobody wrote up
Slack AI Sits where the undocumented knowledge actually is Chat has no canonical answer — it surfaces what was said, not what is true
In-house RAG on SharePoint or Confluence An engineer had a quarter free; looked cheap The failure mode list above, in order. Usually dies on permissions or staleness rather than retrieval quality

Two of these are worth treating as buying signals rather than competition. A disappointing Copilot rollout means budget, executive belief, and a felt failure that was specifically a grounding failure. An abandoned in-house project means the team has already priced the problem and learned the hard part isn't the model. Both cohorts have pre-qualified themselves against discovery question five.

Revised profile, with the market data folded in

Organizations of roughly 200 to 2,000 people — beneath the horizontal platforms' economic floor, above the point where everyone simply knows each other — in permission-sensitive, judgment-dense work where the correct answer changes monthly. Strongest signal: they are currently living with a bundled assistant that disappoints, or the remains of an in-house retrieval project that stalled on permissions or staleness.

Epistemic status. The six characteristics are a reasoning frame argued from how the improvement levers differ — no win/loss data sits behind them. The market section is sourced, but unevenly: revenue and adoption figures come from company announcements and named research, while pricing floors and failure rates come from competitor marketing and secondary posts with no traceable methodology, and are tagged accordingly. Vendor comparison content in this category is written almost entirely by vendors. A dozen of your own closed-won and closed-lost accounts, tested against these six characteristics, would still beat everything on this page.

Sources

  1. Glean surpasses $300M ARR and $200M ARR — company announcements, Dec 2025 / May 2026
  2. Glean business breakdown — Contrary Research
  3. Harvey revenue and funding — Sacra, Jul 2026
  4. RAG vs fine-tuning decision framework — Winder.ai, citing Menlo Ventures' State of Generative AI in the Enterprise
  5. Glean total cost of ownership — GoSearch (a direct Glean competitor; pricing figures are adversarial marketing, treat as indicative only)
  6. Knowledge management tool comparison — Coworker AI (vendor-authored)
  7. Why Copilot agents fail — M365 FM
  8. Enterprise RAG failure modes — dev.to (failure-rate figure unsourced; failure modes corroborated elsewhere)
  9. Enterprise search landscape — Kore.ai (vendor-authored), referencing the Forrester Wave: Cognitive Search Platforms, Q4 2025