MIT PROFESSIONAL EDUCATION
APPLIED AGENTIC AI · CAPSTONE 8.1
Fast for the parts,
right for the person.
The board meeting starts in eight minutes, the executive’s hearing aid battery is nearly depleted, and the charger remains on his kitchen counter.
The appropriate charger arrives discreetly, meeting momentum carries on uninterrupted, and the entire process remains fully confidential and seamless.
This page is the design that removes the dependency. Four of the paper’s seven figures are instruments here rather than pictures: build a request and watch which of the five triggers fires, move the two numbers the business case turns on, read the same request as the audit log keeps it, and walk the three phases to the one that has a date.
01 CONTEXT AND PROBLEM SHEET 02 OF 06
Today the loop runs through me.
I am the North Texas principal lead for executive IT support at an aerospace and defense company. Much of the work is getting bespoke peripherals into people’s hands before they travel, and those requests reach me personally — ten to a hundred a week, by text and walk-up, with more arriving in ways nobody logs.
None of the work is difficult. That is the problem. Every step passes through one person’s availability, so when I travel the requests wait on a time zone, and when the population grows the queue grows with it. My own hours are the only lever left.
Working harder does not fix it: every step passes through one person’s availability. §01 · CONTEXT AND PROBLEM STATEMENT
02 USE CASE AND TECHNOLOGY SHEET 02 OF 06
Five triggers, and never a guess.
The agent takes intake and fulfillment for non-controlled commodity gear and hands a request to a person on any of five triggers. It is not a confidence threshold and it is not a model deciding how sure it feels — each one is a condition you can read off the request.
Build a request below and watch what fires. The lamps are the paper’s figure 02, wired up.
$180
Five conditions, read off the request. No confidence score, no threshold on a model’s certainty — the agent acts on its own only when all five are clear.
All five are clear, so the request goes straight through: entitlement checked, order placed, receipt issued.
A trigger does not reject a request. It goes to a person with everything the agent already gathered attached, so the human starts from the middle of the work rather than the beginning of it.
The agent never refuses when a human could say yes. §02 · THE HANDOFF
Why Microsoft 365 Copilot. Governed, seat-licensed, already in the stack, and inside the tenant boundary with our data. A self-hosted open-weight model costs less per token and gives more control over behavior — and hands me the security posture, the patching and the model lifecycle. In this sector that is a second job, so the tenant boundary decides it for the closed-source platform. The agent holds no credential a person would not hold.
03 COST CONSIDERATIONS SHEET 03 OF 06
The cost is a trip that did not work.
Year 1 is $24k to $42k, almost entirely internal labor: 120 to 200 build hours, platform and security review, and roughly 60 hours of delivery and materials across two tiers. Nothing extra on subscriptions, because Copilot seats are already licensed for everybody in scope. Year 2 onward is $6k to $12k a year for maintenance, monitoring and governance time. The gear itself is not in this number.
Against that sits one failure: an executive travels without the thing that makes the trip work. Airfare and lodging spent, the executive’s hourly rate across the trip hours, and a cost for the meeting that did not happen. The last term does the damage — and it is the one nobody can hand you a number for, which is why the two numbers below are yours to move.
$20,000 — the paper’s estimate
3
The build is fixed at the paper’s $24k–$42k band, midpoint $33k. Move the two things the paper cannot know yet and watch where break-even lands.
The gold band is the paper’s Year 1 build, $24k to $42k. The green bar is what the avoided failures are worth at the numbers you set.
At the paper’s own estimate the build pays for itself somewhere between two and three avoided failures.
This is a projection and the paper says so. The per-failure figure is an estimate, the count is not measured today, and the KPI that carries the money is “avoided trip failures, counted from Crawl”. The pilot has to supply the count before any of it means anything.
04 SECURITY PLAN SHEET 03 OF 06
The record shows what happened without showing what it was for.
Aerospace and defense sets the starting condition: the agent never touches export-controlled technical data. A scope fence holds it to non-controlled gear and a classification check at intake enforces the fence, so it never depends on an assumption.
One thing that fence does not cover is the detail on the cover of this page. The request names a medical device. The agent needs that detail to find the right part — and nothing downstream needs it at all.
Intake holds the whole request, because finding the right charger requires knowing what it charges.
That protects the person’s privacy and their dignity. §04 · IDENTIFY
The log still keeps the request ID, the timestamps, the trigger that fired, the catalog class and the rule that decided — so a GDPR explanation request can be answered without the medical detail ever entering the record.
IT, security, procurement, and me as a standing group.
Three internal data types. Scope fence to non-controlled gear, classification check at intake.
Least privilege on the write path. No standing credential. Five triggers as approval gates.
Receipt with an ID and a clock on every request. Nightly reconciliation against downstream.
I triage operational errors. Data, access, or controlled-item events go to the incident process.
No new system of record. If the agent is pulled, requests fall back to the manual path.
Most of it sits in Identify and Protect. Every record the agent writes downstream is tagged agent-originated, so an auditor can tell agent from person later.
06 SCALING STRATEGY SHEET 04 OF 06
One phase has a date. The other two open on gates.
A roadmap where every phase carries a date is a roadmap that has decided in advance that nothing will go wrong. Here only Crawl is committed — days 0 to 90. Walk and Run are gated, not scheduled: they open when the numbers say so, and a miss holds the next gate closed until the number comes back.
Walk is also where the design stops being supervised and starts being trusted. Crawl puts a human on every request; at Walk the five triggers stop it and everything else runs unattended, which is the whole reason the trigger set has to be right before the gate opens.
Intake only. One pilot group, one item class. The agent captures, validates and routes — a person places every order.
- HUMAN IN THE LOOP
- Every request. Nothing leaves without a human touching it.
- Two consecutive clean weeks at normal volume: no orphaned requests, no missed escalation.
The write path opens. The agent splits into intake, entitlement and fulfillment, and the catalog widens by item class.
- HUMAN IN THE LOOP
- Five triggers stop it. Everything else runs unattended.
- Two more clean weeks with the write path live, reconciliation still returning zero.
Full non-controlled commodity catalog, every supported site, every time zone.
- HUMAN IN THE LOOP
- Exceptions and samples. People see what the agent stopped on, plus a monthly slice of what it did not.
- Monthly review, quarterly scope. Widening scope goes back through the group.
Three data feeds have to be ready before Crawl opens
The catalog needs current part numbers and stock status. The request history needs cleaning, since it doubles as context for the agent and as the replay set for testing. And the entitlement rules have to be written down — that is the hard one: a real share of them exist today only as things I know.
08 KEY PERFORMANCE INDICATORS SHEET 05 OF 06
Three numbers, and each one answers a question.
A wall of metrics gets read once and then ignored. These three stop at three because each one is a question somebody would actually ask, and the money answer lives in section 03 rather than taking a fourth card.
- ASKS
- Does the agent beat a text?
- BASELINE
- 1 to 4 business days. It stretches with my travel.
- TARGET
- Same business day at Run, in any time zone.
- CHECKED
- Monthly
- ASKS
- Do people choose it?
- BASELINE
- Nothing today — every request is a person.
- TARGET
- A rising share resolved end to end without a handoff.
- CHECKED
- Monthly
- ASKS
- Can the write path be trusted?
- BASELINE
- Unknown, because nothing counts them.
- TARGET
- A hard zero. Any orphan is a defect, not a rate.
- CHECKED
- Nightly by the sweep, monthly by the group
Results go to the standing group on the same schedule, and a miss holds the next phase gate closed until the number comes back. The group decides whether the gate or the design was wrong.
The failure mode the numbers would hide
People stop asking overnight and say nothing. No complaint gets filed; the requests simply stop arriving, and containment looks wonderful. So the standing group watches the mix of requests as well as the volume, and a category that goes quiet gets investigated before anybody counts it as a saving.
09 CONCLUSION SHEET 05 OF 06
All of it rests on one choice, and it is not mine.
The design is sound on paper: the platform is already licensed, the boundary holds, the triggers are readable and the phases are gated. None of that matters if the traveler texts me anyway. Every traveler has to choose the agent over texting me — and the only argument that wins is speed. It has to beat sending me a message. That is the whole bar.
However, all of it rests on the people who text me today choosing the agent instead. §00 · EXECUTIVE SUMMARY
The paper
Six sheets, a title block on every one, seven figures. It is drawn as a drawing set rather than written as a report because that is how the work it describes gets checked.