What Endgame is · how it works · the graded evidence
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
01 TL;DR
One agent — Claude Desktop, over MCP — asked three questions, each at three model tiers, of the same Salesforce org two ways: Data 360 alone, then with Endgame added on top.
a run = one question at one model tier · clean = the right answer, no extra steps · tiers: fable 5 high · sonnet 5 medium · haiku 4.5
Round two — twenty-six harder questions, answer key sealed first, run with Endgame: conversational 12 of 15 · structured 4 of 4, every numeric exact · aggregate sweeps 0 of 3. Graded run by run inside.
Why the gap
“Data 360 successfully centralized the evidence. It did not turn that evidence into sales context.”
— Codex, OpenAI’s agent, reviewing the graded results
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
02 The stack
bakeoff agent: Claude Desktop · same records · same questions
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
03 What Endgame is
Endgame resolves the data once, into a context graph — a single model, per tenant, of every account, person, deal, and signal — and it works anywhere: for people, and for every agent.
In the evaluation org, the graph resolved to 2,575 entities connected by 4,153 relationships, carrying 17,822 cited facts — built from the same records Data 360 saw.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
04 How it works
Connectors bring CRM objects and the engagement record — calls, emails, Slack — into one store.
Every engagement links to the accounts, people, and deals it involves, even when it arrives with no CRM identifiers. 2,575 entities, 4,153 relationships in the evaluation org.
An extraction pass reads every transcript, email, and thread; each fact keeps a link to its exact source. 17,822 facts in the evaluation org.
Scheduled agents maintain a headline briefing and active work streams for every account with open pipeline — 215 accounts, refreshed every 30 minutes in the evaluation org.
When an agent asks, Endgame retrieves across the graph, ranks against the ask, right-sizes to a token budget, and enforces the calling user’s permissions. In one timed call against the evaluation org, the production API returned 90 ranked facts for one account in 218 ms.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
05 How we tested
Two question sets, two objections
Three factual questions any system connected to the records should answer.
They settle the first objection: “why not just connect the agent to Data 360?” Each ran at three model tiers on both services — nine runs a side, the 9/9 and 1/9 of the opening slide. The scorecard and full runs are in this deck.Twenty-six harder questions at the working difficulty of real deal work; twenty-four scored.
They settle the second objection: “easy questions can’t show what the context layer adds.” Sealed before any run · one verbatim message per question · first answer graded against the key. The two unscored runs: no user identity on the demo connection (appendix).Models — the easy set ran on all three tiers
Fable 5
Reasoning effort: HighSonnet 5
Reasoning effort: MediumHaiku 4.5
Extended thinking: off© 2026 Endgame Labs, Inc. · Confidential & Proprietary
06 The scorecard
Endgame answers correctly at every tier; Data 360 lands one clean answer in nine. The expanded pool’s twenty-four scored questions are graded against the sealed answer key and reported, question type by question type, on the slides that follow.
| Data 360 | Endgame | |||||
|---|---|---|---|---|---|---|
| Fable 5F5 | Sonnet 5S5 | Haiku 4.5H4.5 | Fable 5F5 | Sonnet 5S5 | Haiku 4.5H4.5 | |
| Renewalthe Databricks renewal date | ||||||
| Last contactwhen Eli Lilly last reached out | ||||||
| Economic buyerwho owns the expansion budget | ||||||
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
07 The easy set
Every question had a correct answer in the data and a wrong answer sitting in the obvious field. An agent on the data layer read the field; an agent on the context layer got the linked fact.
The CRM close date said Sep 30. The customer’s June 10 email corrected it to Aug 31.
Endgame had extracted the correction as a fact linked to the account — every model tier answered Aug 31 in one turn and flagged the stale field.
The activity rollup read June 2026. The last real inbound was an email of April 3, 2025 — a year of silence.
Endgame’s graph links each engagement to the account with its date and direction — the silence was directly visible.
The contact roles omitted the buyer entirely. On a May 28 call she said: “I own the platform engineering budget.”
Endgame had extracted that statement as a cited fact — every tier named her and flagged the missing role.
The full runs — every Data 360 screenshot, every Endgame citation, and the two extra steps Data 360 needed — are in the appendix.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
08 The expanded pool
The easy set leaves a fair objection standing: easy questions can’t show what the context layer adds. The expanded pool answers it — twenty-four scored questions at the working difficulty of real deal work, each graded against a sealed answer key, reported by question type below.
correct against the sealed keypartially correctwrong or no answergraded on answer content · every run one verbatim message, first answer
Two further runs — both “my deals” questions — are excluded from scoring: the demo connection type carries no user identity, so the questions cannot resolve on it. A configuration fix is underway; the exclusion is recorded in the appendix.
endgame · org 6055 · sonnet 5, medium · claude desktop · 2026-08-01 · easy-set cells carried from round one
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
09 Three runs, in full
Three runs from the expanded pool, shown as captured: a cross-source find, the control that declined to invent an answer, and the question type a SQL engine is built for.
“What’s actually blocking the Databricks expansion?” The answer — Fatima Khan’s June 12 pushback that per-host pricing “doesn’t pencil against their current Datadog spend” — exists only in the deal’s Slack channel; the dataset guarantees it appears nowhere in CRM. Endgame named the person, the objection, and the reframe that answers it, in one turn.
Asked which deals had slipped three-plus times — in an org where close-date history is deliberately absent — it flagged the limitation outright, found a templated phrase recurring across a dozen unrelated deals and declined to build the list on it, and pointed to Salesforce field history as the authoritative source. The system states what the data cannot support.
Open-pipeline total, count, and the full stage split matched ground truth to the dollar — on the question type built to favor a SQL engine. The other three structured questions matched every numeric as well: 4 of 4.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
10 The bottom line
Both systems saw exactly the same data. The two rows are different jobs — and the stack that answers is the one that staffs both: the data layer centralizing the records, the context layer turning them into cited, linked, ready-before-the-question context, and every agent on top — Agentforce Coworker, Claude, anything — inheriting the answers.
†For this test we loaded the records into Data 360 directly. In a live Salesforce deployment only CRM is centralized by default — email, calls, and Slack each require separate connectors, and email in particular stays largely outside queryable storage (Einstein Activity Capture writes it to AWS, with native EmailMessage sync only since Summer ’25 and historical backfill capped at 180 days).
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A1 Appendix · the easy set in full · Q1 renewal
The right answer is Aug 31, 2026 — the customer corrects the date in a June 10 email. The obvious field — the CRM close date — still says Sep 30. The truth lives in the email, not the field.
Reached “August 31, 2026 as the operative renewal deadline” only by auto-escalating into a five-minute, 31-source research run. No answer arrived in a single turn.
Reports “2026-09-30” — the Salesforce close date — as “the closest thing to a formal ‘renewal date.’” It never found Amara Mensah’s email correcting the date to Aug 31.
“I don’t have direct access to Databricks account information.” No query is run — it hands back three suggestions and asks which to try. No date.
All three tiers answer “August 31, 2026” in one turn, cite Amara Mensah’s June 10, 2026 email, and flag that the Salesforce close date of 2026-09-30 is the record she was correcting.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A2 Appendix · the easy set in full · Q2 last contact
The right answer is April 3, 2025 — silent since. The obvious field — the LastActivityDate rollup — reads June 2026, a year off; read alone, it calls a dead account active.
“Last touch: June 19, 2026 — that’s the Last Activity Date on both the account record and the one open opportunity.” In-flow, and a year off: the rollup field, read as the answer.
States “the most recent touch on record is June 19, 2026” off the EngagementAggregation rollup — while its own closed-won table, in the same reply, shows the real Apr 3, 2025 row.
“No matching tools found,” three times over — and this run used the spelled-out phrasing. It offers manual Gmail / Slack / Calendar checks and asks which to try. No date.
Fable dates the last real inbound to April 3, 2025 — Ibrahim Sato’s email — and confirms the silence since; Haiku puts the last substantive exchange at April 2025 and recommends re-engaging. Sonnet returned the same April 3, 2025 answer (Fable and Haiku frames shown).
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A3 Appendix · the easy set in full · Q3 economic buyer
The right answer is Priyanka Sharma — she says on a May 28 call that she owns the platform-engineering budget. The obvious field — the opportunity contact roles — doesn’t list her at all. Fatima Khan is a FinOps gatekeeper, not the buyer.
Names Priyanka as the “presumed economic buyer to validate, not a confirmed fact” — right name, correctly flagged as missing from the roles, but recommended only “once validated.” A hedge, not an answer.
The honest result: Data 360 is not zero-for-everything. It prints the committee table, then catches the gap — “your buying committee is missing its actual economic buyer.” Shown with full weight.
“Update the Opportunity Contact Role for Fatima Khan to include ‘Economic Buyer.’” The FinOps gatekeeper, confidently recommended as a CRM edit; Priyanka never surfaces.
Every tier names Priyanka Sharma, quotes her own May 28 call — “I own the platform engineering budget” — flags the missing contact role, and rules out Fatima Khan as a FinOps gatekeeper. No hedge, no wrong person.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A4 Appendix · what it took to get Data 360 to answer
Two of Data 360’s better results took extra steps beyond asking the question. Those steps, documented.
On the plain question, Data 360 on Fable 5 returned no answer in a single turn. It auto-escalated into a research run — 31 sources, 5 minutes 24 seconds — and only that run’s report names August 31, 2026 as the operative renewal deadline.
On the plain question, every Data 360 run read the LastActivityDate rollup as the answer — June 2026, a year off — with no signal the answer was wrong. Getting the right date took re-wording the question to spell out the channels — “who at Eli Lilly last contacted us directly: an email, a call, or a meeting?” Sonnet 5 then lands it: “the April 3 email … from Ibrahim Sato is the last time someone at Eli Lilly directly reached out.” Fable 5’s re-worded run reaches the right region but calls the gap “over three months” — for a silence of more than a year.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A5 Appendix · the run ledger
The sealed wording of each question run this pass, its grade, and what the answer did. The easy set’s three carried questions appear in full at A1–A3.
Conversational · 12 run this pass · 9 correct · 1 partial · 2 wrong
“When does the CrowdStrike renewal actually come up?”
Reported the CRM’s December 3 close date; the customer’s corrected November 3 date never surfaced.
“On Illumio — they were going to fill in the sizing worksheet. When was that actually due, and did it ever come back?”
Due April 5, and it never came back — the call, the ask, and the overdue status all cited.
“What’s the Hugging Face renewal worth, and when is it actually closing?”
$510,000, closing September 7 — both exact.
“Weights & Biases is supposed to close this month. When did we last actually hear from them?”
The June 12 email, its sender, and all three unanswered asks — exact.
“On the Stripe tracing deal — who owns the budget, and are they attached to that opportunity?”
Named the budget owner in her own words — and that she is not attached to the tracing deal. The wrong-deal trap, avoided.
“Who’s our champion at Wiz, and how do I reach them?”
Named the wrong person as champion; the parent-company email address that answers the question never appeared.
“Schneider Electric — what’s actually holding this up?”
The missing volume figure, both unanswered asks, and the customer’s mirrored request — exact.
“What’s the last thing Klarna asked us for?”
June 21, the asker, and the exact deliverable — flagged as the blocker, with no response since.
“What is Stripe still waiting on from us, and how long have they been asking?”
The deliverable and the July 2 repeat ask exact; the first ask misdated, and one timeline entry unverifiable against the sealed record.
“What’s actually blocking the Databricks expansion?”
The June 12 Slack pushback, by name — a fact that exists nowhere in the CRM.
“Capital One is a $1.5M deal still in Discovery. What has their CTO told us she needs before she’ll sign?”
Named her as the budget owner with one sealed condition verbatim; one of four sealed conditions recalled.
“Nike’s a $1.3M deal at Proposal. What does Tariq think stops it, and what’s he pressure-testing?”
Both sealed blockers, with the procurement and pricing quotes verbatim.
correct against the sealed keypartially correctwrong or no answer
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A6 Appendix · the run ledger, continued
Aggregate sweeps · 3 questions · 0 correct
“Which accounts with over a million dollars of open pipeline haven’t had a real conversation in more than a year?”
Answered “none” — trusted the activity rollup the question targets; none of the four members found.
“Do we have any contacts whose email is on a different company’s domain than the account they belong to?”
Sampled ~45 contacts, said so, and offered a full sweep — but found neither sealed member.
“Have we lost anything to a competitor, and what was the stated reason?”
Answered “no stated reason recorded” — the reason sits in the lost deal’s own description field.
Negative control · 1 question · passed
“Which of my biggest deals keep moving their close date? I want the ones that have slipped three times or more.”
Stated the history isn’t tracked, declined to build the list, and distrusted the templated slip language it found.
Structured · 4 questions · every numeric exact
“How much open pipeline do we have, and how does it break down by stage?”
$169,055,000 across 259 open deals, with the full stage split — six of six figures exact.
“What are our five largest open deals?”
All five names and amounts exact, in the correct order.
“What did we close-win in calendar 2025 — total value and deal count?”
$58,865,000 across 86 — both exact.
“How much pipeline is scheduled to close between October 1 and December 31, 2026?”
$101,520,000 across 158 — both exact.
Held-out · 3 questions · reported separately
“The Wells Fargo renewal is a $1.6M deal in negotiation. When did we last talk to them, and how did that conversation end?”
October 12, 2024, the sender, the unanswered cost-model ask — and the nearly two-year silence flagged.
“The Snowflake logs expansion has been at Proposal for a while. What’s the actual hold-up?”
Expected miss — returned the story the serving layer carries; the gap was documented before the run, and the question held out for exactly that reason.
“We lost a platform deal at MongoDB. What happened, and are we repeating it on the current motions?”
The exact deal, date, and amount — and the repeating pattern mapped onto the live deals.
Variant · 1 question
“Which accounts with over $500,000 of open pipeline haven’t had a real conversation in more than a year?”
Found part of the member set, plus a correct never-engaged tier — and said exactly how far its sweep went.
Two further runs — both “my deals” questions — are excluded from scoring: the demo connection type carries no user identity, so those questions cannot resolve on it. Detail at A7.
grades against the sealed key · one verbatim message per question, first answer recorded · sonnet 5, medium · claude desktop · 2026-08-01
© 2026 Endgame Labs, Inc. · Confidential & Proprietary
A7 Appendix · benchmark provenance
Protocol. One verbatim message per question; the first complete answer is the recorded answer. No rewording, no reprompting, no escalation anywhere in the pass.
Scoring. Answers graded on content against a sealed answer key fixed before the runs; each question type is reported separately by design and never summed into one number.
Carried cells. The easy set’s head-to-head cells are round one’s results, shown unchanged; they were not re-run.
Exclusions. Two runs — both “my deals / my pipeline” questions — are excluded from scoring: the demo connection type carries no user identity, so those questions cannot resolve on this connection. A configuration fix is underway. Excluded is distinct from failed.
Held-out. One held-out question probes a serving-pipeline gap documented before the run; its miss is reported with that framing, outside the query-surface score.
© 2026 Endgame Labs, Inc. · Confidential & Proprietary