Positioning note
Of all the ways a company can improve LLM output — evals, harness, context, post-training — which companies should put nearly everything into the context layer, and what that implies for who to sell to.
The tempting split is facts versus skill: facts go in context, skill gets post-trained. That split is wrong, and wrong in the way that matters most — because a great deal of what an organization needs from a model is judgment, and judgment lands on both sides of the line.
"How we decide things here" is judgment, and it belongs in context. "Which of these deals smells like it will churn" is also judgment, and it belongs in weights. Sorting them requires taking judgment apart first.
| Component | Example | Statable? | Home |
|---|---|---|---|
| Criteria — what matters | Concentration risk outweighs growth rate | Usually | Context |
| Thresholds — where the line sits | Discounts over 20% need VP sign-off | Yes, and volatile | Context |
| Weighting — relative importance | Weight retention twice logo count | Sometimes | Either |
| Pattern recognition | This one resembles the accounts that churned | Rarely | Weights |
| Register — how output should feel | House voice, hedging discipline | No, but demonstrable | Weights |
Four of the five are statable. Only pattern recognition genuinely wants weights. That decomposition is why the thesis survives the obvious objection: once you take judgment apart, most of it turns out to be context-shaped.
If a competent owner can write the rule in a sentence, it belongs in context. If the honest answer is "show me five hundred examples and you'll develop a feel for it," you are fitting a function nobody can specify — that is a weights problem.
The volatility test applies to judgment exactly as it applies to facts. Risk appetite shifts quarterly. Bake it into weights and you have frozen last quarter's posture into an artifact you cannot inspect.
If yes, it lives in context regardless of how the first two tests came out. Weights cannot be cited.
Context can emulate weights. Weights cannot emulate context. A rubric plus a handful of examples approximates post-trained judgment — imperfectly and expensively, but it works. Nothing approximates contextual judgment with weights, because what you need is instance-specific and permission-bound.
Context is therefore the general mechanism, and post-training is a compilation step: what you reach for when a piece of judgment is invoked on every call and costs too much to keep re-injecting. It is an optimization, not a knowledge strategy. Read that way, "spend 90% on context" stops being a claim about which lever is more powerful and becomes a claim about which one is load-bearing.
Same lending desk, four pieces of judgment, three different destinations. The question was never facts versus judgment.
Post-training judgment needs a labeled corpus of your own decisions with scored outcomes — thousands of them, ground truth attached. Lenders have it. Insurers, ad platforms and high-volume support organizations have it. Almost nobody else does. Most companies weighing whether to fine-tune their judgment could not do it if they decided to.
Baking what should have been written. Judgment goes into weights; later someone asks why the model said what it said and there is nowhere to look. The judgment becomes unlocatable and unfixable. In regulated work that isn't a performance problem, it's an audit finding.
Injecting what should have been baked. Forty pages of style guide in every prompt. Attention dilutes, adherence drifts, cost scales with every call. This is where post-training legitimately wins, and where context maximalists lose.
Context deserves the overwhelming share of effort when this decomposition keeps landing on the same side. In 2026 it usually does — frontier models are rarely too dumb for enterprise knowledge work. They are uninformed, and increasingly, they are uninformed about how the organization decides rather than merely what it knows.
The sharpest single test. If the correct answer to "what's our position on X" changes monthly, fine-tuning bakes in staleness faster than you can retrain. Ask: what's the half-life of a correct answer here? Weeks means the context layer is the only viable lever. Years means they should fine-tune and stop taking your calls.
Two employees ask the same question and correctly receive different answers based on what they're cleared to know. This rules out post-training outright — weights cannot be need-to-know — and it is the hardest property to bolt on after the fact. Companies with real internal compartmentalization aren't merely good prospects; they're companies for whom the alternatives are structurally unavailable.
If everything already lives in one wiki, competent search wins and a context platform is overkill. The 90% case is knowledge spread across chat, email, documents, meetings, and a handful of people's heads — where the real work is reconciling it, not indexing it.
A long tail of questions, each trivially answerable if you knew the fact. Contrast with novel reasoning work — drug discovery, hard code generation — where capability is the binding constraint and context is merely table stakes.
Both conditions must hold, not either. Some companies are context-bottlenecked but their context is one afternoon of digitization away from solved. That's low value and it churns. You want a bottleneck whose fix compounds.
Fast headcount growth, turnover, reorgs, heavy contractor mix. These companies hemorrhage context structurally, and they already feel it as onboarding pain — which means the budget line exists before you arrive.
The asymmetry is depreciation. Model quality arrives free from the labs, so every dollar of post-training races a curve moving underneath you. Harness is commoditizing into frameworks. Both get cheaper to skip each quarter.
The context layer is the only lever that compounds, stays proprietary, and survives a model swap intact. Spending on the others is renting. Spending on context is owning.
Nobody should actually run 90/0 on context versus evals. You cannot steer context work blind — "we added more documents and it feels better" is how these programs die in month four.
The honest framing is 90% of improvement energy on context, with evals as the instrument rather than a competing investment. Practically: a prospect with no measurement discipline isn't disqualified, but you have to bring the ruler, and that's a real cost in the sales motion worth pricing in.
The profile in one line
Mid-size, knowledge-dense, permission-sensitive organizations — roughly 100 to 2,000 people, growing or churning — where the correct answer changes monthly, differs by who is asking, and currently lives in people rather than systems.
Regulated-adjacent sectors and professional services fit the shape well. So does any company whose product is accumulated judgment.
The frame above is argued. This section is sourced — and it moves two of the six characteristics from plausible to evidenced. Claims are tagged sourced where they trace to primary reporting or named research, and soft where the only available figures come from vendor marketing or unsourced aggregation.
sourcedShare of enterprise LLM deployments using retrieval in production versus those relying primarily on fine-tuning, per Menlo Ventures' State of Generative AI in the Enterprise. The pattern held across consecutive annual surveys. The market has already voted for the context layer — the argument is not whether, it's who should go all-in.
sourcedGlean, as of May 2026 — 15 months after crossing $100M, with 700+ enterprise customers and Fortune 500 logo count roughly doubling year over year. More telling than the headline: 85%+ of customers deploy across five or more departments. Context platforms land horizontally, not as a departmental point solution.
sourcedHarvey, July 2026 — 1,500+ customers, 142,000+ lawyers, roughly half the Am Law 100. The clearest evidence for characteristic 2: the fastest-compounding context businesses are in verticals where truth is permission-bound and judgment-dense, not where it is merely voluminous.
softEnterprise retrieval deployments reported as failing or materially underdelivering in year one. Treat the number as directional — it circulates through secondary posts without a traceable methodology. The failure modes are the durable part, and they are consistent across every account: stale source content, ignored permission models, misconfigured connectors. Characteristics 1 and 2 are not theory; they are the top two reasons these projects die.
Glean began by selling to technology companies in the 500–2,000 employee range, then moved decisively upmarket. What it left behind is the relevant fact for anyone entering now.
softReported entry economics sit near a 100-seat minimum and roughly $60K annual contract floor, with all-in cost of ownership for mid-to-large deployments quoted between $350K and $480K per year. These figures come largely from competitor marketing content rather than Glean's own disclosures, so treat the magnitudes as indicative and the direction as reliable: the category leader has a floor, and it sits above most of the mid-market.
The result is a barbell. Large enterprises buy a horizontal context platform. Everyone below the floor assembles something out of wiki tools, bundled assistants, and in-house retrieval projects — which is precisely the cohort the six characteristics describe.
Any prospect in this band already has something. The question in discovery is never "do you have a knowledge tool" — it's which of these they've settled for, and which failure they've already felt.
| What they run | Why they chose it | Where it breaks |
|---|---|---|
| Microsoft 365 Copilot | Bundled, no procurement, already licensed | Grounding against SharePoint and permissions is the most-cited reason its agents disappoint; users expect a consumer chatbot and get an assistant that doesn't know the company |
| Glean, Coveo, Sinequa | Genuine cross-system breadth; the only category that doesn't force an ecosystem choice | Seat minimums and total cost put them out of reach below roughly 1,000 employees |
| Confluence + Rovo | AI bundled into every paid tier at a few dollars per user; already the engineering wiki | Only as good as what someone wrote down, and stays largely inside Atlassian's own surfaces |
| Guru | Named expert re-verifies each card on a schedule — the most conservative answer to "is this still true" | Human curation doesn't scale to a wide answer surface; knowledge has to fit the card shape |
| Notion AI | Zero friction if the company already lives in Notion | Knows Notion. Does not know Slack, email, the CRM, or the meeting nobody wrote up |
| Slack AI | Sits where the undocumented knowledge actually is | Chat has no canonical answer — it surfaces what was said, not what is true |
| In-house RAG on SharePoint or Confluence | An engineer had a quarter free; looked cheap | The failure mode list above, in order. Usually dies on permissions or staleness rather than retrieval quality |
Two of these are worth treating as buying signals rather than competition. A disappointing Copilot rollout means budget, executive belief, and a felt failure that was specifically a grounding failure. An abandoned in-house project means the team has already priced the problem and learned the hard part isn't the model. Both cohorts have pre-qualified themselves against discovery question five.
Revised profile, with the market data folded in
Organizations of roughly 200 to 2,000 people — beneath the horizontal platforms' economic floor, above the point where everyone simply knows each other — in permission-sensitive, judgment-dense work where the correct answer changes monthly. Strongest signal: they are currently living with a bundled assistant that disappoints, or the remains of an in-house retrieval project that stalled on permissions or staleness.
Epistemic status. The six characteristics are a reasoning frame argued from how the improvement levers differ — no win/loss data sits behind them. The market section is sourced, but unevenly: revenue and adoption figures come from company announcements and named research, while pricing floors and failure rates come from competitor marketing and secondary posts with no traceable methodology, and are tagged accordingly. Vendor comparison content in this category is written almost entirely by vendors. A dozen of your own closed-won and closed-lost accounts, tested against these six characteristics, would still beat everything on this page.