Case study · written procedure & AI accuracy

The page was worth twenty points

In April 2026 a benchmark put three frontier models against 99 questions over a retail database, twice each: once with the schema alone, once with the schema plus a four-kilobyte markdown file a person had written about what the data means. The file was worth 17 to 23 points of accuracy. Switching models was worth nothing. This page is about that file — and about why the file is the thing most companies adopting AI have not written.

Aug 27, 20264 sources

The demo takes four seconds. An analyst types a question into a chat box wired to the company warehouse — revenue last quarter, net of returns, by channel — and the model writes the SQL, runs it, and puts a number on the screen. The number is specific. It has decimals. The room is impressed. Nobody present can say whether it is right.

About half the time, it is not. Not absurdly wrong — wrong the way a smart new hire is wrong in week two. The model summed gross sales because nobody told it returns live in a second table. It counted orders where the question meant net revenue. It picked the wrong fact table from three plausible candidates. Each of those mistakes is a decision a senior analyst makes correctly, silently, dozens of times a day, without noticing a decision was made.

In April 2026, two researchers at Cube — a company that sells a semantic layer, which is worth keeping in view — measured exactly this. Rumiantsau and Fokeev ran 99 natural-language questions over the Cleaned Contoso Retail dataset in ClickHouse. Single shot, medium reasoning effort, no tools, no retries. Three models: Claude Opus 4.7, Claude Sonnet 4.6, GPT-5.4. Each ran twice. The first condition got the database schema. The second got the schema plus one hand-written markdown file — 8,969 characters, about 2,200 tokens — covering which fact table to use when, how each measure is calculated, the hierarchies, the data quirks, the defaults, and what the ambiguous words mean. With the schema alone, the three models passed between 45.5 and 50.5 percent of the questions. With the file, 67.7 to 68.7 percent.

Six configurations, two bands

The finding

The model did not matter. The page did. Pass rate on 99 questions, each dot one model-condition pair, bars the Wilson 95% intervals, sorted by pass rate.

4050607080%Schema onlySchema + one written page+17 to +23 ppGPT-5.468.7Claude Sonnet 4.668.7Claude Opus 4.767.7Claude Opus 4.750.5Claude Sonnet 4.646.5GPT-5.445.5
Schema onlySchema + the 4 KB document

Source: Rumiantsau & Fokeev, Cube, arXiv 2604.25149 (April 28, 2026), n = 99, single-shot. Answers graded by Claude Sonnet 4.6 — a judge that shares a model family with two of the three systems it graded.

Look at where the intervals sit. Every comparison across the two clusters is significant at p < 0.01. Every comparison inside a cluster is noise — p ≥ 0.42 among the schema-only runs, p ≥ 0.79 among the runs with the document. Moving between three of the most capable models money can buy changed nothing a significance test could see. Four kilobytes of prose moved every model by at least 17 points.

The paper’s own explanation is the right one: the document converts open-ended inference into constrained lookup. Without it, the model must guess which of several fact tables holds the truth about sales, guess whether margin means the column or the formula, guess how this company counts a customer. With it, those questions are answered before they are asked. The file’s contents are not documentation in the shelfware sense. They are the twenty-odd decisions a senior analyst carries in their head and applies from memory, question by question — written down once, with the ambiguity removed, each with an answer and an owner. A procedure for operating on this company’s data. There is an older, plainer name for a page like that, and every operations team already uses it.

The document is not magic. Roughly 30 percent of questions still failed with it, and the failures concentrate where prose stops helping: percentiles, correlation, standard deviation, ABC classification, threshold ranking, SQL that needs recursion or several chained CTEs. A page fixes ambiguity. It does not fix arithmetic.

Cube’s paper also collects a short literature of the same experiment run elsewhere, and the pattern across it deserves its own chart. These rows are secondary — reported via the Cube paper, not re-verified here — but they all lean the same way.

The floor sets the gain

Literature

Accuracy before and after adding written context, six studies, sorted by where each started. The less written down going in, the more a written page is worth.

0255075100%Bayer †pharmacovigilance · GPT-4+70.08.378.3data.world †199-table schema · GPT-4+37.516.754.2BIRD †one evidence sentence · GPT-4+20.034.954.9CubeContoso · three models+17 to +2345.5–50.567.7–68.7dbt Labs †modeled project · GPT-5.3-Codex+15.984.1100.0dbt Labs †modeled project · Sonnet 4.6+8.290.098.2
Without written contextWith written context

Source: studies marked † as cited by the Cube paper (arXiv 2604.25149), secondary; the Cube row is that paper’s own result, shown at its range across three models.

Bayer’s pharmacovigilance queries started at 8.3 percent — a jargon-dense domain with almost nothing written down — and a context document took GPT-4 to 78.3. dbt Labs started at 84.1 percent on a fully modeled project, where most of the decisions were already encoded, and the same move was still worth 15.9 points. Two negative results sharpen the mechanism. In the Bayer study, narrowing the schema to the relevant tables without the document only cut the failure rate to 50 percent; the document did the rest. And buried in Cube’s own paper: an agentic setup with tools, where the model could browse the equivalent semantic knowledge in a catalog, did not beat the schema-only baseline. Having the page somewhere the model could look it up was worth roughly nothing. Putting the page in the prompt was worth twenty points.

What the page cost

Cost

Extra input tokens the document added per question, against the accuracy it bought. The page never exceeded a fifth of the prompt. It carried all of the gain.

ModelTokens the page addedPoints it boughtLatency
Claude Opus 4.7+4,199+17.2 pp+0.6 s
Claude Sonnet 4.6+2,939+22.2 pp+1.8 s
GPT-5.4+2,361+23.2 pp−0.9 s

Source: arXiv 2604.25149, input-token and latency columns; deltas are doc-condition minus schema-only.

This is the part that generalizes, because the mechanism has nothing to do with SQL. A runbook is the same object: the twenty decisions an on-call engineer makes from memory at 2 a.m., written down with the ambiguity removed. So is a deployment procedure. So is a definition of done. An agent pointed at a wiki and told to search is the catalog condition — the one that failed to beat the baseline. An agent handed the page is the document condition. A lot of AI adoption work underway right now is building the catalog condition and calling it context.

There is a second version of this finding hiding in the season’s tooling news, told from the other side. On August 18, 2026, Snowflake’s Semantic View Autopilot began ingesting Power BI semantic models directly; Tableau workbooks have made the same trip since at least January, and Databricks’ Genie Code has done the equivalent import into Unity Catalog metric views since June. Snowflake published a support table saying exactly what carries over — I wrote about the migration itself in Signal issue #16. Laid out as a strip, the support table makes the same argument as the benchmark.

What survives the trip

Categorical

Elements of a BI semantic model, laid out from what a warehouse ingestion keeps to what it leaves behind. This strip is categorical — read off a vendor support table, not measured.

Survives

  • Tables
  • Columns
  • Relationships
  • Column descriptions
  • Metric definitions

Partially ports

  • Report-level measures

Stays behind

  • Time intelligence (PREVIOUSMONTH, SAMEPERIODLASTYEAR, TOTALYTD)
  • Tableau Level of Detail calculations
written in plain SQLwritten in the tool’s dialect

Source: Snowflake Semantic View Autopilot support documentation (Power BI ingestion GA August 18, 2026; Tableau since at least January); Databricks Genie Code, Unity Catalog metric views (June 2026).

Read the strip left to right. Tables, columns, relationships, descriptions, metric definitions — the things describable in plain SQL — travel to the next system. Report-level measures make part of the trip. Time intelligence does not: PREVIOUSMONTH, SAMEPERIODLASTYEAR, TOTALYTD stay where they were written. Tableau’s Level of Detail calculations stay too. Databricks is candid about the residue — its documentation recommends attaching a screenshot of the original dashboard so the agent can check its numbers against the picture. The half of a system written in a portable language can be handed to the next reader. The half written in the tool’s dialect cannot, and someone re-derives it by hand. The next reader, increasingly, is a machine.

The practical move fits in a sentence. Write the four-kilobyte file for your own domain: twenty decisions, in prose, each with the question it answers, the answer, and the person who owns it. Which table is the source of truth. What the fiscal calendar means. What counts as a customer, a conversion, a done. Do not automate it first, and do not buy a tool that promises to generate it — the document in this study worked because a person who understood the data wrote down what they knew, and 8,969 characters turned out to be enough room.

The three models in this benchmark will look dated within a year, and the leaderboard order among their successors will keep shuffling. The page describing what Contoso’s tables mean will still be true.

Sources: Rumiantsau & Fokeev, arXiv 2604.25149 (April 28, 2026) and the Cube semantic-layer benchmark repository; literature band as cited therein (Bayer / JAMIA Open 2024; data.world / Sequeda 2023; BIRD 2023; dbt Labs 2026); Snowflake Semantic View Autopilot support documentation; Databricks Genie Code documentation (June 2026).