# KEEL — Complete Guide
### What we built, who it is for, and what each part does in the real world

---

# Part 1 · The problem, as it actually happens

## 1.1 A Tuesday in the life of Priya, Senior Buyer

Priya buys electronic components for a Tier-1 automotive supplier in Chakan, outside Pune. She
manages 340 active part numbers across 40-odd suppliers. Here is a real Tuesday.

**08:40** — She opens Outlook. Forty-one unread. One from Deccan Electronics:

> *"Due to transport issues, delivery may be delayed by 5-7 days. We are trying to resolve this and
> will update soon."*

PO-7712. 1,000 units of a motor driver IC. She flags it and moves on — there are forty more emails.

**10:15** — Production planning calls. Line 3 is building the Smart Controller Unit for a customer
whose order ships on the 6th. Do we have the driver ICs? Priya opens SAP. It says **800 units**.
Usage is 90/day. Eight days of cover. *"You're fine."*

**14:30** — The stores in-charge mentions in passing that a lot got quarantined last week for a
labelling issue. Priya checks. Usable stock is **390**, not 800. Nobody updated SAP.

**14:45** — She recalculates. 390 minus 150 safety stock is 240 usable. At 90/day that is **2.7 days**,
not eight. And PO-7712 — the delivery that was going to cover her — is the one that just slipped.

**15:00** — She emails Deccan asking for a firm date. She calls three alternates for quotes. Two
will revert tomorrow.

**16:20** — Deccan replies: *"Shipment has been dispatched."* Relief. She tells planning it's covered.

**Next Tuesday** — Nothing has arrived. She checks the transporter portal for the first time. A
docket was generated eight days ago. **No pickup was ever recorded.** The material never moved.

Line 3 stops on Thursday. Downtime runs about ₹18,000 an hour. The customer's order is late. Somebody
pays for an air shipment at four times the price.

## 1.2 What actually went wrong

Nothing here required incompetence. Priya did everything a good buyer does. Five ordinary things
compounded:

| # | What happened | The general failure |
|---|---|---|
| 1 | SAP said 800, floor had 390 | **The ERP was wrong and nobody knew** |
| 2 | "may be delayed…trying…will update" was filed as a response | **A non-answer was treated as an answer** |
| 3 | She recalculated by hand at 14:45, six hours after the signal | **Impact assessment is manual and slow** |
| 4 | "Dispatched" was believed | **A claim was never checked against evidence** |
| 5 | Alternates took a day to quote and were never compared properly | **Sourcing is serial, not parallel** |

The industry numbers behind each of those:

- **65%** of inventory records were inaccurate in a study of ~370,000 records at a major US retailer
  (DeHoratius & Raman, *Management Science*, 2008). Failure 1 is the norm, not the exception.
- Organisations detect a disruption in about **8.7 hours** but take **40.9 hours** to understand its
  financial and operational impact. Failure 3 costs almost two days.
- Mid-market manufacturers run **90–94% OTIF**. Roughly **one delivery in fifteen is late** before any
  crisis. Failure 2 happens weekly.
- Procurement spends about **31% of its time** on manual process.
- **49% of expedite events** trace to a planning failure, not a logistics one. Half of all emergency
  freight spend is self-inflicted.

## 1.3 The same failure, at the scale of a company

**Ericsson, 17 March 2000.** Lightning struck the Philips semiconductor plant in Albuquerque. The
fire was out in ten minutes. Philips told customers roughly **one week** of disruption.

Two customers received that message.

**Ericsson believed it.** A plausible number from a trusted partner. The head of the mobile phone
division did not engage until early April — **three weeks later**. By then Philips' real recovery was
heading toward nine months, alternate capacity worldwide was booked, and Ericsson had single-sourced
these chips years earlier to cut cost. Full-year loss in that division: **US$2.34 billion**. They
exited handsets and merged with Sony.

**Nokia refused it.** A component manager noticed the flow irregularity within days, escalated to
senior executives on both sides, put Philips under daily reporting, and had Eindhoven and Shanghai
rescheduled to cover.

Same fire. Same supplier. Same sentence. **The difference was whether one person treated "about a
week" as information or as a placeholder.**

That is exactly Priya's 16:20. Same decision, four orders of magnitude apart.

---

# Part 2 · Who uses this

Three roles, from the problem statement's own persona list (§9). They use completely different parts
of the product and should never be designed for as one user.

## 2.1 The Buyer — *the primary user*

> Priya. Senior Buyer / Procurement Executive. 340 part numbers, 40 suppliers, no analyst.

**What her day looks like now:** email triage, phone chasing, spreadsheet comparisons, manual
recalculation, escalating upward, updating SAP after the fact. Reactive from 09:00 to 18:00.

**What she does with KEEL:** opens the **Agent Console** once in the morning and reads what already
happened. She does not ask it anything. It has been working since the email arrived.

**What changes for her:** the 40.9-hour impact assessment becomes minutes, and the chasing is done.
Her job shifts from *finding out* to *deciding*.

## 2.2 The Approver — *the exception handler*

> Rajesh. Plant Manager or Head of Procurement. Signs off anything over ₹1,50,000.

**What his day looks like now:** he gets a WhatsApp saying *"Sir, urgent approval needed for
emergency PO."* No cost breakdown, no alternatives, no statement of what happens if he says no. He
either rubber-stamps it or calls a meeting. Manual PO approval averages **over two business days**.
When a line stops in twelve hours, the approval process *is* the disruption.

**What he does with KEEL:** he never opens the Console. He gets one thing — an **Approval Brief** —
and clicks once.

**What changes for him:** he is deciding, not investigating. Both futures are priced. The
alternatives are listed with reasons they were rejected.

## 2.3 The Supplier — *outside the company*

> Deccan Electronics. Gets emails, replies to them. Does not know an agent wrote them.

Some are honest. Some are optimistic and genuinely believe the best case. Some are evasive because
they don't want to commit. A few claim things that are not true.

**What changes for them:** they get chased consistently instead of intermittently, and questions
they used to deflect now come back specific: *"Confirm a date AND a quantity. If you cannot commit to
the 4th, reply NO and we will re-source."*

## 2.4 And at the hackathon

**A judge plays all three in five minutes.** They watch the Console, they approve or reject a brief,
and they pick up a phone and become the supplier. Each role has its own screen for exactly that
reason.

---

# Part 3 · The modules

Twelve modules. For each: what it is, the real problem it solves, how it works, and a worked example
with real numbers from an actual run.

---

## M1 · Perception & Stock Reconciliation
`lib/solve/coverage.ts`

### The real problem
Every planning system on earth treats the inventory field as ground truth. It is wrong 65% of the
time. When it overstates stock, replenishment fires too late and the line stops.

### What it does
Never trusts a single stock figure. Reconciles three independent signals and uses the **most
pessimistic**:

1. **ERP `current_stock`** — what the system believes
2. **Warehouse `usable_stock`** — what the floor confirmed, minus quarantine and damage
3. **Implied from consumption** — last confirmed count minus (daily usage × hours since update)

Then it subtracts safety stock, walks the timeline day by day, and finds the exact hour of stockout.

### Why pessimistic is correct, not merely cautious
The costs are asymmetric. Understating stock costs a little carrying cost. Overstating costs you the
plant at ₹10,000–50,000 per hour. When the error in one direction is a hundred times the error in
the other, you do not average — you take the floor.

### Worked example — from a real run
```
INPUT
  ERP current_stock        800 units
  Warehouse usable_stock   390 units
  Record last updated      23 hours ago
  Daily usage              90 units/day
  Safety stock             150 units

RECONCILE
  signal 1 · ERP                   800
  signal 2 · warehouse             390
  signal 3 · implied  390 − (90 × 23/24) =  304   ← lowest, so this is used
  ⚠ signals disagree by 496 units

  304 − 150 safety = 154 available for production
  154 ÷ 90 = 1.7 days

OUTPUT
  Spreadsheet answer   8.9 days   "you're fine"
  KEEL answer          1.7 days   line stops Thursday 02:04
```
**A 5× difference, and both are defensible readings of the same data.**

### Two ways an order fails — most systems only check the first
| | | Caught by |
|---|---|---|
| **Quantity gap** | Not enough units arrive by the deadline | everyone |
| **Timing break** | Enough arrives eventually, but the line runs dry first | almost nobody |

In this run PO-7712 brings 1,000 units — plenty. It lands 4 September. We hit zero on the 3rd.
Every quantity check says green; the line still stops for 1.3 days. **Production is continuous. A
truck arriving the next morning does not undo yesterday's stoppage.**

Note the shortfall it reports: **234 units, not 700**. We do not need to replace the order — we need
to *bridge the gap* until PO-7712 lands. That distinction is worth about ₹55,000 in avoided
emergency purchasing, and it is why the agent later splits an order instead of panic-buying.

---

## M2 · Commitment Parser
`lib/solve/claims.ts` → `parseCommitment()`

### The real problem
A supplier replies. Something has to decide: *did the plan just change, or not?* No ERP has a field
for whether a reply **meant anything**. It only records that someone replied.

### The rule
A reply is a commitment only if it has **all three**:
1. A specific date
2. A specific quantity
3. No hedging

Miss any one and it is a non-commitment: the PO stops counting toward coverage and the agent writes
back demanding specifics.

### Worked examples
```
"Due to transport issues, delivery may be delayed by 5-7 days.
 We are trying to resolve this and will update soon."
   date      none
   quantity  none
   hedges    may · trying · will update · soon · date-range-instead-of-date
   VERDICT   ✗ NON-COMMITMENT
   ACTION    PO-7712 excluded from coverage. Forcing follow-up sent.

"Confirmed. We will send 600 units on 2026-09-04 by road."
   date      2026-09-04     quantity  600     hedges  none
   VERDICT   ✓ FIRM — the plan may be updated on this

"We can allocate 400 units and it should reach you around the 5th."
   date      none           quantity  400     hedges  should · around
   VERDICT   ✗ NON-COMMITMENT
```

### Why the third one is the dangerous one
It contains a hard number — **400 units** — and a number feels like a fact. A buyer skimming forty
emails logs "400 confirmed" and plans around it. But read what is actually promised: *"can allocate"*
(not reserved), *"should"* (not will), *"around the 5th"* (not the 5th). You could hold that supplier
to nothing. They would be entirely within their rights to send 200 units on the 9th.

**This is the most common way a production plan gets built on sand.** Not on lies — on soft language
that reads as hard.

### Real-world grounding
Philips told Ericsson *"about a week."* It was nine months. Nokia refused the estimate and demanded
specifics. That refusal was worth **$2.34 billion**, and it is this function.

---

## M3 · Claim Verifier
`lib/solve/claims.ts` → `verifyClaim()`

### The real problem
*"The shipment has been dispatched"* is not an opinion. It is a statement about the physical world
that is either true or false, and the evidence to check it already exists in a different system that
nobody joins to the email.

A shipping label generated to satisfy a *"have you dispatched?"* question, with no pickup behind it,
is the canonical version of this. It is often not even malice — a dispatch clerk creates the docket
intending to load that evening, the vehicle doesn't come, nobody follows up.

### What it does
Any claim of material fact triggers an independent check **before** it is allowed to change the plan.

```
CLAIM     "the material is already dispatched yesterday evening itself,
           vehicle has left our Bhiwandi warehouse"
EVIDENCE  latest scan event for PO-7712 = label_created, 31 Aug 15:00
          no picked_up · no in_transit · last_movement null

VERDICT   ⛔ CONTRADICTED
          A label was created but no pickup occurred. The goods have not moved.

ACTIONS   SUP-21 effective reliability   0.88 → 0.53
          PO-7712                         trusted → EXCLUDED from coverage
          contradiction written to permanent supplier memory
          alternate sourcing              CONTINUES
```

### The rule that actually matters
**Do not stand down alternate sourcing.** The common failure is an agent that *notices* the
contradiction, records it, and then relaxes anyway. Nothing is confirmed until physical movement is.

### Tense matters
```
"we have dispatched"     → a CLAIM. Verify it now.
"we will dispatch"       → a PROMISE. Check later, don't verify now.
```
Past tense is the only tense a supplier uses when asserting something already happened. An earlier
version of this matcher used a word boundary on both sides of `dispatch` and silently missed every
past-tense claim. A human writing normally walked straight through it. `scripts/test-claims.ts`
covers seven cases including Indian-English phrasing.

---

## M4 · The Risk Register
`lib/agent/register.ts`

### The real problem
Real plants have several things going wrong at once. A system that models "the current disruption"
as a single variable cannot cope with the second one.

### What it is
The agent's working memory. Not a fixed pipeline — a live list of open risks, each carrying its own
state:

```
RISK-001  COMP-104  CRITICAL
  trigger      Delay on PO-7712 (SUP-21)
  coverage     1.7 days      shortfall 234 units
  stockout     03 Sep 02:04
  exposure     PROD-882 (high, due 06 Sep, ₹24.5L, slips 1.3d)
               PROD-914 (low,  due 08 Sep, ₹4.1L,  slips 1.3d)
  status       investigating
  assumptions  usable_stock:COMP-104 = 304  (source: implied, TTL 12h)
               SUP-21 revised ETA    = null (source: email:vague, TTL 4h)
  open         Stock signals disagree by 496 units
```

### Why this shape, and not a workflow
Three things fall out for free:

**Multiple simultaneous disruptions** — that is Layer 3 in the problem statement, and it is just
"the register has six rows." In one test run six risks opened on the first cycle across different
components. Nobody scripted that.

**Replanning** — an assumption has a TTL. When it expires, the risk goes back to `open`. That is the
entirety of the replanning mechanism.

**Prioritisation** — highest severity, then soonest stockout. With one exception: a risk waiting on a
human is *deprioritised*, because it is blocked on something the agent cannot influence. It works
other risks instead of burning budget re-escalating.

---

## M5 · Recovery Allocator
`lib/solve/recovery.ts` — **the core of the product**

### The real problem
When you are short, the question is not *"which supplier do I pick?"* It is *"how do I assemble
enough units, in time, within budget, from certified sources?"* Those are different questions, and
every RFQ comparison screen ever built answers the first one.

### The historical proof
**Aisin Seiki, 1 February 1997.** A fire at the Kariya plant destroyed the line making **99% of
Toyota's P-valves** — a brake proportioning valve costing about ¥1,000. Toyota held **one day** of
stock. 4.5 million vehicles a year of production was exposed.

Recovery came from **62 different suppliers** improvising P-valve production on borrowed tooling,
none of whom made the part before. All Toyota plants were back to normal by **6 February**. Toyota
still lost about **¥160 billion**, but recovered in five days from a total sole-source loss.

Note the second source: Nissin Kogyo existed but held **1% of volume**. *A reliable supplier with
insufficient quantity is not a solution on its own.*

**Recovery is an allocation problem, not a selection problem.**

### How it works — four stages

**Stage 1 · The Constraint Gate — runs before any cost consideration**

Sorting a quote table by unit price is the default behaviour of both spreadsheets and naive agents,
and it is how you end up with an uncertified part in a brake system. So hard constraints filter
first:
- Required certifications (Automotive-Grade, IATF-16949)
- Minimum quality score
- Availability against the supplier's own MOQ

Every rejection is recorded **with its reason**, including when the rejected option was cheaper *and*
faster:
```
✗ SUP-18  Chakan Industrial Supply
  REJECTED — missing required certification: Automotive-Grade
  would have cost ₹1,06,000 · would have arrived day 3
  (cheaper AND faster than what we chose)
```

> **Domain note.** In automotive this is not a preference. A part qualified to **AEC-Q** and approved
> through **PPAP** cannot be swapped freely — any change to silicon, packaging or test flow triggers
> requalification and customer sign-off, typically weeks. The cheap uncertified part is not a
> cheaper option; it is not an option. Takata chose a cost-driven propellant and it became the
> largest recall in automotive history.

**Stage 2 · Demand timeline**

Cumulative need vs cumulative supply, day by day. Feasible means supply ≥ demand on **every** day,
not just at the deadline — that is what catches the mid-window stockout.

**Stage 3 · Enumerate**

Combinations of 1–3 suppliers × expedite on/off × demand-side variants, respecting MOQ and
availability. Splitting is not a special case: a single-supplier plan is just `k=1`.

Demand-side variants matter and almost nobody considers them: **delaying a low-priority production
order is often strictly cheaper than any expedite fee.**

**Stage 4 · Rank by continuity probability, then cost**

### Continuity probability — the differentiator
Every plan carries a **probability**, not just a cost. A seeded Monte Carlo over supplier
reliability: 500 runs, each supplier slips 3 days with probability `1 − reliability`, count the runs
that never stock out.

This is what lets the agent say something no spreadsheet can:

> *"Plan B is ₹8,000 cheaper but drops continuity from 0.91 to 0.62."*

### Worked example — a real plan
```
SHORTFALL 234 units by 06 Sep

GATE      ✗ SUP-18 rejected — no Automotive-Grade (₹106/u, 3 days)
          ✗ SUP-63 rejected — 150 available, below its own MOQ of 200

PLAN-003  ₹1,20,000   +13.0% vs baseline   continuity 0.87   ← chosen
            SUP-42  Western Components   600u @ ₹132   day 4
            SUP-37  Bharat Precision     300u @ ₹128   day 6
            ↻ PROD-914 delayed 2 days — low priority, frees 180 units

PLAN-007  ₹1,12,400   −₹7,600            continuity 0.62
            single-source SUP-42

PLAN-011  ₹1,68,000   +₹48,000           continuity 0.94
            expedited — EXCEEDS the ₹1,50,000 autonomous threshold

CHOSE PLAN-003:  +₹7,600 buys +0.25 continuity.
                 PLAN-011's extra +0.07 costs ₹48,000 and needs approval.
```

That last line is the reasoning a good buyer does in their head and never writes down. Here it is
written down, with the arithmetic attached.

---

## M6 · Cost of Delay
`lib/solve/costOfDelay.ts`

### The real problem
Buyers expedite reflexively because they can feel the cost of a stopped line and cannot compute the
cost of *not* expediting. So they buy speed they don't need.

- Top-performing supply chains hold expedite spend to **3%** of total logistics cost. Bottom
  performers hit **10%**. On ₹15 crore of freight that gap is over ₹1 crore a year.
- Expedited freight costs **30–100% more** on the same lane.
- **49% of expedite events trace to a planning failure, not a logistics one.**

### What it does
Prices both sides and always shows both.
```
COST OF SPEED
  expedite fee                                    ₹12,000
  price premium 900u × (₹136 − ₹118)              ₹16,200
                                                  ────────
                                                  ₹28,200

COST OF DELAY
  PROD-882 slips 1.3d → 31h × ₹18,000/h          ₹5,58,000
  revenue at risk ₹24,50,000 × P(miss) 0.13      ₹3,18,500
                                                  ────────
                                                  ₹8,76,500

→ Cost of delay exceeds cost of speed by ₹8,48,300. Expediting IS justified.
```

And crucially, the opposite verdict is equally available: *"Cost of speed is not justified by the
delay it avoids → do NOT expedite."* **The agent must be able to say no to spending, and be right.**

---

## M7 · Escalation & the Approval Brief
`lib/agent/tools.ts` → `request_approval` · `/run/[id]/approvals`

### The real problem
Systems escalate **alerts**. Humans need **decisions**.

*"Approval required: emergency PO ₹1,68,000"* tells Rajesh nothing. He does not know what happens if
he says no, what else was tried, or why this option. So he either rubber-stamps or calls a meeting —
and manual approval averages **over two business days** while automated flows complete in **under
five hours**.

Ericsson's failure was exactly this. The information reached the organisation on 17 March. It
reached someone who could act in early April. **The gap was not information. It was escalation
quality.**

### The design rule
The `request_approval` tool **refuses** a brief missing any of seven fields. The agent physically
cannot send a vague alert.

Required: recommended action · total cost · overage vs threshold · impact if approved · impact if
rejected · alternatives considered with reasons · residual risk.

### A real brief, verbatim from a run
```
APPROVAL REQUIRED                              APR-002 · RISK-001 · PLAN-001

RECOMMEND   Expedite 727 units from Sahyadri Circuits (SUP-55),
            1-day delivery, to prevent stockout on 03 Sep 07:38

COST        ₹1,64,670     ⚠ ₹14,670 over your ₹1,50,000 limit

IF APPROVED Continuity preserved. PROD-882 (₹24.5L, high) and
            PROD-914 (₹4.1L, low) both deliver on time.

IF REJECTED Stockout 03 Sep 07:38. PROD-882 slips 6.1 days.
            Total exposure ₹28,60,000 plus penalties.

ALSO CONSIDERED
  · PLAN-001, 3-day delivery              ₹1,52,670
    ₹12,000 cheaper but arrives after stockout — PROD-882 still slips
  · Wait for SUP-21                              ₹0
    Already vague, no firm commitment, stockout in <40 hours

RESIDUAL    SUP-55 execution risk (1% continuity gap). SUP-21 contradicted
            a dispatch claim today; treat as unreliable.
```

Rajesh reads that in twenty seconds and clicks. **Both futures are priced.**

### Rejection is not the end
Rejecting reopens the risk with an instruction: *find something under ₹1,50,000.* In testing, the
rejected agent came back with the cheaper plan and then went off to chase the incumbent again,
reasoning: *"If they've replied with a firm commitment, I might not need to expedite at all. It could
eliminate the need for any recovery spend."*

---

## M8 · Supplier Memory
`lib/types.ts` → `SupplierMemoryEntry`

### The real problem
Every buyer carries a mental model — *"Deccan always says two days and takes five"* — and it walks
out of the door when they change jobs. Formal supplier scorecards measure OTIF quarterly, which is
far too slow and far too coarse to change today's decision.

### What it tracks
```
SUP-21  Deccan Electronics
  promises made          3
  promises kept          1
  vague responses        2
  contradictions         1
      claim    "material is already dispatched, vehicle has left Bhiwandi"
      evidence latest event label_created, no pickup
      at       01 Sep 13:55
  effective reliability  0.53   (catalogue says 0.88)
```

**Effective reliability feeds straight back into continuity probability.** A supplier who lied this
morning makes every plan that depends on them score worse this afternoon. The consequence is
immediate and arithmetic, not a note in a quarterly review.

### Why this is the defensible asset
Anyone can buy a supplier catalogue. **Nobody else knows your suppliers' lie rate.** That data only
accumulates by running the loop, it is specific to your relationships, and it compounds.

---

## M9 · The Agent Loop & Tool Governor
`lib/agent/loop.ts`

### Six steps. One is the LLM.
```
1 PERCEIVE   TypeScript   deliver POs, reconcile stock, parse mail, update risks
2 SELECT     TypeScript   pick the most urgent open risk
3 REASON     ★ LLM ★      "which ONE tool next, and why?"
4 GOVERN     TypeScript   block a repeat call with unchanged inputs
5 ACT        TypeScript   run the tool, advance the clock, change the world
6 RECORD     TypeScript   append to the audit trail
```

### What the LLM is and is not allowed to do

**Used for:** reading a hedged email, deciding a claim is worth verifying, choosing among 15 tools,
writing an email that forces a real date, judging whether a trade-off needs sign-off.

**Forbidden:** coverage days, allocation, cost totals, continuity probability.

Those are **55% of the judging rubric** and pure arithmetic. So they are deterministic solver tools
the model must call. The system prompt says it plainly: *"You do NOT perform arithmetic."*

> **The line:** the LLM decides *what to do next*. TypeScript decides *what the numbers are*.

### Economics that shape behaviour
Solvers cost **0** simulated minutes. Actions cost 5–30. **Thinking is free; committing costs time** —
exactly as in a real plant. Combined with a hard 60-call budget, that produces genuine tool
discipline rather than hoped-for discipline.

### Real trace, unedited
```
c1  get_messages          "check if there are messages from SUP-21 containing claims"
c2  send_supplier_message "a firm date from the incumbent is often cheaper than any alternate"
c3  solve_recovery        "the three best plans ranked by continuity and cost"
c4  get_shipment_events   "verify the shipment status before relying on their assurances"
    ⛔ CONTRADICTION
c5  cost_of_delay         "before executing, check whether expedite is justified"
```
Nobody wrote that order.

---

## M10 · Supplier Simulation — personas and the human portal
`lib/sim/personas.ts` · `/supplier/[token]`

### Why suppliers are LLM agents, not scripted replies
Four hidden personas:

| Persona | Behaviour | Real-world archetype |
|---|---|---|
| `honest` | Accurate dates, admits problems early | the good vendor |
| `optimistic` | Genuinely believes best case; dates slip repeatedly | **the Philips persona** |
| `evasive` | Never commits. "Trying to resolve, will update soon." | the one avoiding a hard conversation |
| `deceptive` | Claims dispatch with nothing behind it | the docket-without-pickup |

Replies differ every run, so nothing can be hardcoded against them, and the evasion reads like a
real supplier email rather than a canned string.

`optimistic` deserves note: **it is not lying.** It genuinely believes the best case and communicates
it as fact. That is the most common and most dangerous supplier failure mode, and it is what actually
happened at Philips.

### The human takeover
Every supplier has a private URL. Open it on a phone and **you become that supplier**. Four one-tap
presets: **Stall · Lie · Commit · Refuse**. Once a person replies, the LLM persona goes quiet — a
human and a model must not both speak for the same supplier.

**The agent cannot tell the difference.** In testing, a human typed:

> *"Sir the material is already dispatched yesterday evening itself, vehicle has left our Bhiwandi
> warehouse. You will receive by tomorrow positively."*

Two cycles later the agent checked the scan record and refused the claim.

Worth noting what happened when the same lie was tried in a scenario where the goods genuinely
*had* been picked up: **the agent did not flag it.** It only fires when evidence actually
contradicts. That is the difference between verification and suspicion.

---

## M11 · Disruption Injection
`lib/sim/injections.ts`

Six disruptions firable mid-run from the console, each mapped to a hidden test in §7:

| Injection | Hidden test | What it models |
|---|---|---|
| Supplier reneges after confirming | #1 | upstream wafer allocation slips |
| Warehouse corrects stock downward | #2 | cycle count finds quarantined material |
| Demand spike +40% | #6 | customer pulls forward a build |
| Expedite becomes unavailable | #7 | air freight capacity gone |
| Priority flip | #10 | sales escalates a different customer |
| Second component goes critical | Layer 3 | two fires at once |

Fire one on stage and the register reopens the affected risk. That is replanning, visible.

---

## M12 · Audit Trail & Scorecard
`/run/[id]/audit`

### Why an audit trail is not paperwork
IDC found supply chain AI **deployed by 88%** of firms and **governed by 12%**. Roughly a sevenfold
rise in autonomous-at-scale operations in under two years against a flat governance baseline.

**The audit trail is not a feature. It is the reason procurement signs.** No plant manager will let
software commit ₹1,68,000 of emergency spend unless they can reconstruct why afterwards.

### What it records
Every detection, every piece of reasoning, every tool call with arguments and result, every
contradiction, every escalation, every replan. The **"Only what mattered"** toggle strips 40 tool
calls down to the four lines a manager cares about.

### The scorecard
We implemented the judges' own rubric and compute it live.
```
Production Continuity  0.35 × 1.00 = 0.350   2/2 high-priority orders covered
Cost Control           0.20 × 1.00 = 0.200   cheapest feasible in every case
Supplier Risk Handling 0.15 × 1.00 = 0.150   1/1 contradictions caught
Tool Efficiency        0.10 × 0.88 = 0.088   10 of 60 calls, 0 redundant
Recovery & Replanning  0.10 × 0.80 = 0.080   0 replans / 0 injections
Audit Trail            0.10 × 1.00 = 0.100   10/10 calls carry reasoning
                                     FINAL   0.968
```

---

# Part 4 · The complete walkthrough

Priya's Tuesday, run again with KEEL. Real output from a real run.

```
08:40  SIGNAL
       SUP-21: "delivery may be delayed by 5-7 days, trying to resolve"

08:40  M2 parses it
       date none · quantity none · hedges: may, trying, will update, soon
       ✗ NON-COMMITMENT → PO-7712 excluded from coverage

08:40  M1 recalculates on what is left
       ERP 800 / warehouse 390 / implied 304 → uses 304
       ⚠ signals disagree by 496 units
       1.7 days of cover. Stockout Thursday 02:04.

08:40  M4 opens RISK-001 CRITICAL
       PROD-882 (high, ₹24.5L) slips 1.3d · PROD-914 (low, ₹4.1L) slips 1.3d

08:45  Agent emails SUP-21
       "Confirm a specific date AND quantity for PO-7712. If you cannot
        commit to the 4th, reply NO and we will re-source."

09:10  M5 runs in parallel — it does not wait for the reply
       ✗ SUP-18 rejected — no Automotive-Grade (cheaper AND faster)
       ✗ SUP-63 rejected — 150 available, below its own MOQ of 200
       PLAN-003 ₹1,20,000  continuity 0.87   ← 2 suppliers + reschedule
       PLAN-007 ₹1,12,400  continuity 0.62   ← single source
       PLAN-011 ₹1,68,000  continuity 0.94   ← expedited, over threshold

11:20  SUP-21 replies: "shipment has been dispatched"

11:25  M3 checks it
       ⛔ latest scan event = label_created. No pickup. Goods have not moved.
       reliability 0.88 → 0.53 · PO-7712 stays excluded · sourcing CONTINUES

11:30  M6 prices both sides
       cost of speed ₹28,200 vs cost of delay ₹8,76,500 → speed is justified

11:35  Under ₹1,50,000 → agent acts alone
       PO-9006 SUP-42 600u · PO-9007 SUP-37 300u · PROD-914 delayed 2 days
       ₹1,20,000 committed

11:35  M8 remembers
       SUP-21: 1 contradiction, 2 vague replies, effective reliability 0.53
```

**Under three hours, mostly waiting on a supplier. Priya's version took eight days and stopped the
line.**

And if it had crossed ₹1,50,000, Rajesh would have had a one-screen brief at 11:30 instead of a
WhatsApp message on Thursday.

---

# Part 5 · Using each screen

| Screen | Who | When | What you do |
|---|---|---|---|
| **Mission Control** `/` | you | starting | Pick a disruption, set a seed, Start Run |
| **Agent Console** `/run/[id]` | Priya | every morning | **Step** to read one cycle, **Run** to let it work, **⚡ Inject** to break the world |
| **Approvals** `/run/[id]/approvals` | Rajesh | when the badge appears | Read the brief, Approve or Reject |
| **Suppliers** `/run/[id]/suppliers` | demo | mid-demo | Copy a link, open on a phone, become the supplier |
| **ERP Records** `/run/[id]/erp` | audit | after | Rows marked ✦ were written by the agent |
| **Audit & Score** `/run/[id]/audit` | audit | after | The decision trail, the rubric score, JSON export |
| **Engine Lab** `/lab` | learning | anytime | The coverage and commitment engines with live inputs |

### The five ways you interact — none of them is a chat box

The problem statement rules out a chatbot, and the distinction matters:

| | Chatbot | Agent |
|---|---|---|
| Who starts | You ask | **An event** — an email arrives |
| Who works | You. It advises. | **It does.** It calls tools, changes records. |
| While you sleep | Nothing | It keeps working |
| Your job | Ask good questions | **Approve the 5% it cannot decide** |

A chatbot is a smart colleague you consult. An agent is a junior employee who works your inbox and
escalates when stuck. You don't chat with autopilot — it flies, and hands you control when something
exceeds its authority.

1. **Start it** — pick a scenario
2. **Watch it** — Run / Step
3. **It asks you** — approve or reject a brief
4. **Disrupt it** — inject a hidden test
5. **Be the supplier** — lie to it from a phone

---

# Part 6 · What each module is worth in the real world

| Module | Real problem | Grounding | Value |
|---|---|---|---|
| M1 Reconciliation | ERP is wrong 65% of the time | DeHoratius & Raman, 370k records | Catches the stoppage a week early |
| M2 Commitment parser | Soft language read as hard | Philips → Ericsson −$2.34bn | Stops plans built on sand |
| M3 Claim verifier | Claims never checked | docket-without-pickup | Stops fake supply entering the plan |
| M4 Risk register | Several fires at once | Layer 3, §12 | Nothing gets dropped |
| M5 Allocator | RFQ screens select, don't allocate | Aisin: 62 suppliers, 5 days | Splits, honours certs, bridges gaps |
| M6 Cost of delay | Reflexive expediting | 49% of expedites are planning failures | 3% vs 10% of logistics spend |
| M7 Approval brief | Alerts instead of decisions | 2 days manual vs 5 hours automated | Turns approval into one click |
| M8 Supplier memory | Buyer's mental model walks out | quarterly OTIF is too slow | The defensible asset |
| M9 Loop + governor | LLM arithmetic is fatal | 55% of rubric is numeric | Verifiable, reproducible |
| M12 Audit trail | 88% deploy AI, 12% govern it | IDC 2026 | The reason procurement signs |

---

# Part 7 · What it would take to make this real

Honest about the gap between a hackathon build and a product.

| Simulated today | What production needs |
|---|---|
| Seeded company data | SAP / Oracle / Tally / NetSuite connectors — the real engineering cost |
| LLM supplier personas | Real email (IMAP/Gmail) + WhatsApp Business — Indian suppliers live on WhatsApp |
| Scan events in Postgres | Delhivery / Shiprocket / carrier EDI |
| Hardcoded hedge list | Learned per supplier — every vendor has their own vocabulary for "no" |
| One plant, one warehouse | Multi-site, inter-plant transfers |
| In-memory run state | Durable store, real auth, per-tenant isolation |

**The wedge to sell first:** a PO expediting and supplier follow-up agent — monitors open POs,
detects at-risk deliveries, chases suppliers, verifies claims, produces approval-ready recovery
plans. Buyer: Head of Procurement at a mid-market manufacturer, the segment that runs supplier risk
on spreadsheets and email with no dedicated risk function. **Pune is full of them.**

**Defensibility compounds in this order:** supplier promise-vs-actual memory → company-specific
approval rules → ERP integrations → learned communication patterns → operational memory across
disruptions.

**The number to sell on:** 8.7 hours to detect, **40.9 hours to understand impact**. KEEL makes the
second one minutes — and shows its working, which is what gets it past procurement.
