Issue #17 · August 24, 2026

DoorDash open-sourced the robot that writes their code. I made it do my chores

A couple weeks ago DoorDash released as part of their CLI repo something quietly remarkable: the internal tool their own engineers use to have AI write software. It’s called the agentic-orchestrator, it’s free on GitHub, and most people scrolled past it. I didn’t. I installed it, pointed it at a small project, and spent two days watching how a forty-billion-dollar company actually thinks about AI. This was not in a press release, but in working code.

Here’s what the tool does, in plain terms. You describe a feature you want built. The orchestrator researches your codebase, writes a plan, has AI critics attack the plan (one for architecture, one for security, one for testing), implements the code in a quarantined copy of your project, reviews its own work, and then stops at the door marked ship it — because a human holds that key, always.

I gave it a deliberately small job: add discount codes to a toy shopping-cart module. Three things happened that taught me more than testing random X account ‘success stories’ ever could.

First, it refused to overcomplicate. The planning stage looked at my little feature and ruled it too small to break into phases. One plan, one pass. An AI that says “this doesn’t need more process” is rarer than an engineer or analyst who says it.

Second, it stopped and waited for me. Midway through, it wanted to run a shell command — a file copy, roughly. Instead of granting itself permission, it queued a request for a human and went silent. For thirteen hours, because I wasn’t watching. When I finally answered, it offered to remember my approval as a narrow, scoped rule — trust widened one pattern at a time, never all at once. The override flag exists, but they named it --dangerously-skip-permissions. It hurts to type. That’s the point.

Third, it showed receipts. The final review wasn’t one AI skimming the code, it was three, each with a separate job: one for repo hygiene, one for code quality, one for correctness. The correctness reviewer didn’t read the test suite and nod; it re-ran the tests itself and saved the runner’s raw output to an evidence folder — then went further and invented its own edge-case probes (what happens with a null discount? a zero-value coupon?) and logged those results too. A separate checker verifies the evidence files exist and aren’t padded duplicates before any agent is allowed to declare success. Every prompt, response, and judgment call is written to disk, including a spec ambiguity it hit mid-plan (percent discounts can produce fractional cents), where it recorded the question, its chosen answer, and its own confidence score: 0.60. Underneath all of it, I still didn’t take anyone’s word: I went into the quarantined copy and ran the suite myself. 12 for 12. Total cost of the whole run: $3.01.

Three dollars and one permission slip. That’s what a researched, planned, implemented, reviewed, and tested feature cost.

The deeper pattern is the one worth carrying out of this. DoorDash ships a consumer CLI that lets AI agents search restaurants, build carts, and price orders, everything except pay. The purchase stays human. Their engineering tool hands AI everything except ship. The merge stays human. Two different teams, same bet: let the agents do the work, keep a person’s hand on anything that can’t be undone.

I run my own fleet of scheduled AI agents at home, they watch my calendar, my infrastructure, my writing pipeline, my Strava stats, and I’d arrived at the same rule independently before reading a line of their code: my agents propose, my finger makes the call.

When strangers keep converging on the same law, it isn’t coincidence. It’s an evolution with a visible fossil record. First we engineered prompts — one instruction, one answer. Then we engineered loops — agents that try, check, and try again. Now we’re engineering graphs: webs of agents that plan, criticize, and verify each other. The machines took the middle. Humans kept the ends. The tools are free. The lesson is too.

Crafting runs weekly. The job is to be useful, not to sell. If a line in here is close to a move you're working on, the email reaches a person.

Hiring rather than hunting? Signal is the column for your side of the desk.

Email meRead more Crafting

← All Crafting issues · Drafted with Claude · Edited by Paul Brown