◈ MISSION CONTROL · PORTFOLIO

MIT PROFESSIONAL EDUCATION
APPLIED AGENTIC AI · CAPSTONE 8.1

8:52 A.M. the message lands

Fast for the parts,
right for the person.

The board meeting starts in eight minutes, the executive’s hearing aid battery is nearly depleted, and the charger remains on his kitchen counter.

9:01 A.M.

The appropriate charger arrives discreetly, meeting momentum carries on uninterrupted, and the entire process remains fully confidential and seamless.

This page is the design that removes the dependency. Four of the paper’s seven figures are instruments here rather than pictures: build a request and watch which of the five triggers fires, move the two numbers the business case turns on, read the same request as the audit log keeps it, and walk the three phases to the one that has a date.

01 CONTEXT AND PROBLEM SHEET 02 OF 06

Today the loop runs through me.

I am the North Texas principal lead for executive IT support at an aerospace and defense company. Much of the work is getting bespoke peripherals into people’s hands before they travel, and those requests reach me personally — ten to a hundred a week, by text and walk-up, with more arriving in ways nobody logs.

None of the work is difficult. That is the problem. Every step passes through one person’s availability, so when I travel the requests wait on a time zone, and when the population grows the queue grows with it. My own hours are the only lever left.

FIG 01 THE LOOP THE SAME REQUEST BOTH WAYS
01
Somebody texts Usually the day before they fly.
02
I work out what they need Since it is often not what they asked for.
03
I check entitlement Part of the answer is in a system, part in policy, and part in my memory.
04
I place the order or open the ticket Watch it move, and tell them when it ships.
05
When something goes sideways I fix it Most of the time they never hear about it.
06
The gear is in their hands Nobody else in the room needs to know anything happened.
Working harder does not fix it: every step passes through one person’s availability. §01 · CONTEXT AND PROBLEM STATEMENT

02 USE CASE AND TECHNOLOGY SHEET 02 OF 06

Five triggers, and never a guess.

The agent takes intake and fulfillment for non-controlled commodity gear and hands a request to a person on any of five triggers. It is not a confidence threshold and it is not a model deciding how sure it feels — each one is a condition you can read off the request.

Build a request below and watch what fires. The lamps are the paper’s figure 02, wired up.

FIG 02 HUMAN IN THE LOOP THE FIVE TRIGGERS THAT HAND A REQUEST TO A PERSON

$180

Five conditions, read off the request. No confidence score, no threshold on a model’s certainty — the agent acts on its own only when all five are clear.

1 · MORE THAN ONE Two or more items, or a full kit.
2 · OVER THE LINE Estimated value above roughly $500 cumulative over 30 days, so an order cannot be split to stay under it. A person signs.
3 · TOO FAST Several requests from one person in a short window. Something is wrong or something is urgent. Both need a person.
4 · ADDRESS MISMATCH Ship-to does not match what the backend holds. Catching a wrong address before it ships costs far less than after.
5 · NOT STANDARD Custom, non-catalog, or anything the rules do not cover. Outside the catalog it is a judgment call.
THE AGENT ACTS ON ITS OWN

All five are clear, so the request goes straight through: entitlement checked, order placed, receipt issued.

A trigger does not reject a request. It goes to a person with everything the agent already gathered attached, so the human starts from the middle of the work rather than the beginning of it.

The agent never refuses when a human could say yes. §02 · THE HANDOFF

Why Microsoft 365 Copilot. Governed, seat-licensed, already in the stack, and inside the tenant boundary with our data. A self-hosted open-weight model costs less per token and gives more control over behavior — and hands me the security posture, the patching and the model lifecycle. In this sector that is a second job, so the tenant boundary decides it for the closed-source platform. The agent holds no credential a person would not hold.

03 COST CONSIDERATIONS SHEET 03 OF 06

The cost is a trip that did not work.

Year 1 is $24k to $42k, almost entirely internal labor: 120 to 200 build hours, platform and security review, and roughly 60 hours of delivery and materials across two tiers. Nothing extra on subscriptions, because Copilot seats are already licensed for everybody in scope. Year 2 onward is $6k to $12k a year for maintenance, monitoring and governance time. The gear itself is not in this number.

Against that sits one failure: an executive travels without the thing that makes the trip work. Airfare and lodging spent, the executive’s hourly rate across the trip hours, and a cost for the meeting that did not happen. The last term does the damage — and it is the one nobody can hand you a number for, which is why the two numbers below are yours to move.

FIG 03 PAYBACK AVOIDED-FAILURE VALUE AGAINST YEAR 1 COST

$20,000 — the paper’s estimate

3

The build is fixed at the paper’s $24k–$42k band, midpoint $33k. Move the two things the paper cannot know yet and watch where break-even lands.

The gold band is the paper’s Year 1 build, $24k to $42k. The green bar is what the avoided failures are worth at the numbers you set.

AVOIDED-FAILURE VALUE $60,000 Year 1, at the numbers you set
YEAR 1 BUILD $24k–$42k Midpoint $33k. Internal labor, not gear
BREAK-EVEN 2 failures Against the $33k midpoint

At the paper’s own estimate the build pays for itself somewhere between two and three avoided failures.

This is a projection and the paper says so. The per-failure figure is an estimate, the count is not measured today, and the KPI that carries the money is “avoided trip failures, counted from Crawl”. The pilot has to supply the count before any of it means anything.

04 SECURITY PLAN SHEET 03 OF 06

The record shows what happened without showing what it was for.

Aerospace and defense sets the starting condition: the agent never touches export-controlled technical data. A scope fence holds it to non-controlled gear and a classification check at intake enforces the fence, so it never depends on an assumption.

One thing that fence does not cover is the detail on the cover of this page. The request names a medical device. The agent needs that detail to find the right part — and nothing downstream needs it at all.

FIG 05 THE REDACTION THE SAME REQUEST, TWICE

Intake holds the whole request, because finding the right charger requires knowing what it charges.

That protects the person’s privacy and their dignity. §04 · IDENTIFY

The log still keeps the request ID, the timestamps, the trigger that fired, the catalog class and the rule that decided — so a GDPR explanation request can be answered without the medical detail ever entering the record.

FIG 04 SECURITY THE CONTROLS ACROSS THE SIX NIST CSF 2.0 FUNCTIONS
GOVERN

IT, security, procurement, and me as a standing group.

IDENTIFY

Three internal data types. Scope fence to non-controlled gear, classification check at intake.

PROTECT

Least privilege on the write path. No standing credential. Five triggers as approval gates.

DETECT

Receipt with an ID and a clock on every request. Nightly reconciliation against downstream.

RESPOND

I triage operational errors. Data, access, or controlled-item events go to the incident process.

RECOVER

No new system of record. If the agent is pulled, requests fall back to the manual path.

Most of it sits in Identify and Protect. Every record the agent writes downstream is tagged agent-originated, so an auditor can tell agent from person later.

06 SCALING STRATEGY SHEET 04 OF 06

One phase has a date. The other two open on gates.

A roadmap where every phase carries a date is a roadmap that has decided in advance that nothing will go wrong. Here only Crawl is committed — days 0 to 90. Walk and Run are gated, not scheduled: they open when the numbers say so, and a miss holds the next gate closed until the number comes back.

Walk is also where the design stops being supervised and starts being trusted. Crawl puts a human on every request; at Walk the five triggers stop it and everything else runs unattended, which is the whole reason the trigger set has to be right before the gate opens.

FIG 06 SCALING THREE PHASES · CRAWL HAS A DATE
CRAWL COMMITTED · DAYS 0 TO 90

Intake only. One pilot group, one item class. The agent captures, validates and routes — a person places every order.

HUMAN IN THE LOOP
Every request. Nothing leaves without a human touching it.
GATE TO THE NEXT PHASE
Two consecutive clean weeks at normal volume: no orphaned requests, no missed escalation.
WALK GATED · MONTHS 4 TO 8

The write path opens. The agent splits into intake, entitlement and fulfillment, and the catalog widens by item class.

HUMAN IN THE LOOP
Five triggers stop it. Everything else runs unattended.
GATE TO THE NEXT PHASE
Two more clean weeks with the write path live, reconciliation still returning zero.
RUN GATED · MONTH 9 ON

Full non-controlled commodity catalog, every supported site, every time zone.

HUMAN IN THE LOOP
Exceptions and samples. People see what the agent stopped on, plus a monthly slice of what it did not.
STANDING GATE
Monthly review, quarterly scope. Widening scope goes back through the group.

Three data feeds have to be ready before Crawl opens

The catalog needs current part numbers and stock status. The request history needs cleaning, since it doubles as context for the agent and as the replay set for testing. And the entitlement rules have to be written down — that is the hard one: a real share of them exist today only as things I know.

08 KEY PERFORMANCE INDICATORS SHEET 05 OF 06

Three numbers, and each one answers a question.

A wall of metrics gets read once and then ignored. These three stop at three because each one is a question somebody would actually ask, and the money answer lives in section 03 rather than taking a fourth card.

FIG 07 KPIs BASELINES, TARGETS, AND HOW OFTEN WE CHECK THEM
TIME TO SERVE
ASKS
Does the agent beat a text?
BASELINE
1 to 4 business days. It stretches with my travel.
TARGET
Same business day at Run, in any time zone.
CHECKED
Monthly
CONTAINMENT
ASKS
Do people choose it?
BASELINE
Nothing today — every request is a person.
TARGET
A rising share resolved end to end without a handoff.
CHECKED
Monthly
SILENT FAILURES
ASKS
Can the write path be trusted?
BASELINE
Unknown, because nothing counts them.
TARGET
A hard zero. Any orphan is a defect, not a rate.
CHECKED
Nightly by the sweep, monthly by the group

Results go to the standing group on the same schedule, and a miss holds the next phase gate closed until the number comes back. The group decides whether the gate or the design was wrong.

The failure mode the numbers would hide

People stop asking overnight and say nothing. No complaint gets filed; the requests simply stop arriving, and containment looks wonderful. So the standing group watches the mix of requests as well as the volume, and a category that goes quiet gets investigated before anybody counts it as a saving.

09 CONCLUSION SHEET 05 OF 06

All of it rests on one choice, and it is not mine.

The design is sound on paper: the platform is already licensed, the boundary holds, the triggers are readable and the phases are gated. None of that matters if the traveler texts me anyway. Every traveler has to choose the agent over texting me — and the only argument that wins is speed. It has to beat sending me a message. That is the whole bar.

However, all of it rests on the people who text me today choosing the agent instead. §00 · EXECUTIVE SUMMARY

The paper

Six sheets, a title block on every one, seven figures. It is drawn as a drawing set rather than written as a report because that is how the work it describes gets checked.