Issue #23 · September 27, 2026

The Agent Answers The Question You Modeled For. Vision Is Modeling For The One Nobody Has Asked Yet

The intent did not port

The Four Skills, part 3 of 4 | Vision

A series on the four skills that matter most for analysts now that frontier models and open-source tooling do the tool-native half of the job. Part 1 showed the warehouses ingesting every measure with equal weight. Part 2 was the morning the copy and the original disagreed, and whose job it was to say which one was right. This one is vision: who the model is for, and for how long. Attention, decisions, vision, systems.

“Which model did it read?”

Someone in finance is going to ask you that after a demo goes well, not before it. On September 10, OpenAI turned on a Data agent inside ChatGPT Work that reads whatever semantic layer your organization already has and answers from it. Whatever you modeled, it reads. Whatever you didn’t, it never sees. It answers anyway. This issue is not about that agent. It is about the analyst who decided, a year before the agent existed, what it would find when it arrived.

Every measure in your semantic layer exists because someone asked a question in 2021, or 2023, or last quarter. The warehouse kept the arithmetic. It did not keep the intent — the meeting, the objection, the reason the filter excludes what it excludes. A human reading the dashboard could still stop you in the hallway and ask what a number meant. An agent reads the field’s documentation, or it reads nothing, and it cannot tell you which one it just did.

Vision is the skill of modeling for the reader who hasn’t shown up yet. Not the analyst’s own question, which is answered by Friday, but the head of sales asking about a territory that doesn’t exist until the map is redrawn in Q2. The customer service rep who needs one number at 4:45 on a Friday and will trust whatever the screen says. The agent that has no hallway and no hand to raise. An analyst with vision writes the definition, the boundary, and the reason — in prose a stranger can read — because the stranger is coming, and the stranger might be software.

Here is what that looks like in my own shop. We run BigQuery under Looker, and open-source dashboards are moving in beside Looker because they are cheap to stand up and cheaper to throw away. The vision I am building toward is an agent that sits between the warehouse and every one of those surfaces and reads one thing: our dbt-governed definitions. One metric, written once. Tested. Documented with the reason for every filter. Looker reads it. The open-source dashboards read it. The agent reads it. When the definition changes, it changes in one file, the tests run, lineage tells every downstream surface it moved, and nobody sends an email. The surfaces become disposable because the meaning doesn’t live in them.

That architecture is not a modeling choice. It is a vision choice. It only works if someone sat down before the first dashboard and asked how three different readers would use the number — the VP who wants a trend, the rep who wants a threshold, the agent who wants a definition it can quote — and wrote for all three. Forethought is the deliverable. The YAML is the receipt.

What’s actually moving in the market

The reader I just described is now shipping. OpenAI’s Data agent connects to Redshift, BigQuery, Databricks, Snowflake, and more on one side, builds dashboards in Looker’s competitors on the other, and pulls its understanding of the business from “semantic layers and trusted sources such as Databricks Genie Ontology, dbt, GitHub, Snowflake Horizon, and BI dashboards.” It is not a new modeling standard. It is a stress test of every modeling standard already in production, run by a vendor with no stake in which one wins.

The standard those layers are supposed to converge on is still unfinished. Open Semantic Interchange — the format Snowflake, Databricks, Salesforce, dbt Labs, and more than 50 organizations (up from 17 launch partners) are building together — entered the Apache Incubator on July 10 as Apache Ossie. Converters for dbt’s Semantic Layer and Apache Polaris are merged. The agents arbitrating your definitions are already live. The format meant to keep them honest across vendors is still pre-release.

dbt Labs’ own bridge is narrower and has been shipping longer: an MCP server, live since April 2025 and still getting releases as of July 2026, that exposes a dbt project’s Semantic Layer, lineage, and freshness metadata to any AI client through the same protocol Claude and Cursor use to reach outside data. It doesn’t invent a definition. It hands over the one you already wrote, in a shape an agent can hold. That is the middle piece of the architecture above, and it already exists.

The counter-read: none of this asks you to model differently for the agent’s sake. It asks whether the model you already shipped can survive a reader who can’t raise a hand and ask what you meant.

I watch a smaller version of this every week. My own newsletter fleet reads a task’s notes field the same way an agent reads a semantic layer — it only knows what the last person writing to that field bothered to spell out, and it never asks.

What I’d do this week

User moment: Someone in ops asks an agent — OpenAI’s, or whatever your org already runs — a question your dashboard was never built to answer, and gets a confident number back anyway.

Shape: Take the ten metrics from Part 2’s scoring exercise — owned, tested, documented, tied to a decision. For each, name the reader most likely to be surprised by it four quarters out: a VP after a reorg, a rep after a comp-plan change, an agent after nobody’s around. Write the question that reader will ask. Then check whether the metric’s current documentation, the prose a stranger’s agent would read rather than the SQL, answers it or silently assumes it away.

Time budget: Ninety minutes — ten metrics, roughly nine minutes each once Part 2’s list exists.

Artifact: A one-line addendum per metric: the reader, the future question, and whether the current definition covers it, partially covers it, or goes silent.

Success: At least half of the ten metrics have a documented answer to their own future question before an agent has to improvise one live, in a room you’re not in.

Credible failure: Every metric passes today’s test and fails the four-quarters-out one — proof the scoring exercise measured whether the model works, not whether it keeps working for a reader who hasn’t been invented yet.

Part 4 is systems.

Sources

Crafting runs weekly. The job is to be useful, not to sell. If a line in here is close to a move you're working on, the email reaches a person.

Hiring rather than hunting? Signal is the column for your side of the desk.

Email meRead more Crafting

← All Crafting issues · Drafted with Claude · Edited by Paul Brown