← All workflows

Workflow · August 28, 2026

Build a Manager's AI-Replacement Post-Mortem: Paste a Workforce-Automation Plan, Get a Risk-Rated Lessons Report

✓ TestedFounderFor Founder & Operator
Time saved3-4 hours per workforce-automation review

The task

You're a founder or COO staring at a workforce-automation plan — your own draft, a consultant's deck, or a competitor's leaked memo — and you need to know which assumptions will hold and which will blow up in six months. This piece turns any pasted plan into a risk-rated post-mortem before you sign the org chart. Use it whenever someone in the room says "AI can do that headcount's job."

Before AI

Normally this is a two-week loop: you circulate the plan, pull a few advisors onto calls, hunt for prior-art disasters, and try to write down what you learned in a Google Doc that no one reads. Most founders skip it and just run the experiment. That's exactly how Meta ended up cutting roughly 10% of its workforce in May and then quietly killing the November wave when the AI productivity gains didn't materialize.

The workflow below compresses that loop to under an hour. You paste one artifact, get three outputs: an extracted assumption list, a risk score per assumption, and a lessons report you can hand to your leadership team.

The workflow

1. Paste the plan and extract the load-bearing assumptions

The goal of step one is to force the plan's implicit claims out into the open. Most workforce-automation plans hide their assumptions inside verbs like "streamline" or "consolidate." You want them as testable statements.

Prompt
You are a skeptical Chief of Staff reviewing a workforce-automation plan for a founder-CEO. Read the plan below and extract every load-bearing assumption — the claims that, if wrong, would sink the plan.

Return a numbered table with these columns:
- # (assumption number)
- Assumption (one sentence, in the plan's own terms)
- Type (Productivity / Quality / Adoption / Org-design / Financial / Timeline)
- Evidence cited in the plan (quote or "none")
- Falsifiable? (Yes/No — could a 90-day pilot prove or disprove this?)

Extract 8-15 assumptions. Do not editorialize yet. Do not suggest fixes. Just surface what the plan is betting on.

PLAN:
Sample input
INTERNAL MEMO — CONFIDENTIAL
Project Northstar: AI-Native Reorg, FY26
Author: Sample Founder, CEO, Fabricated Widgets Inc. (staff 420)

Summary. We will reorganize Fabricated Widgets into an AI-native operating model in two waves — Wave 1 in April, Wave 2 in October. Target end-state headcount: 260 (a 38% reduction), with the remaining team operating as "talent-dense pods" of 4-6 humans supervising fleets of AI agents.

Rationale. Our internal coding-assistant rollout in Q4 showed a 3.1x increase in pull requests per engineer. Extrapolating that gain across Support, Marketing Ops, Sales Development, and QA suggests we can hit the same output with roughly 60% of current headcount. Two consulting partners have validated the model.

Wave 1 (April). Eliminate 90 roles across Support Tier 1 (35), SDR (28), Marketing Ops (15), and Junior QA (12). Replace with a stack of three agent frameworks plus a supervisor tool built by our Platform team (est. 8-week build). Expected annual savings: $14.2M.

Wave 2 (October). Eliminate an additional 70 roles across Mid-Market AE, Recruiting Coordinator, FP&A analyst, and Content roles, contingent on Wave 1 hitting >85% of pre-reorg output KPIs by end of Q3.

Governance. A "human-in-the-loop" review layer will catch agent errors for the first 60 days post-Wave-1, then sunset. Quality will be measured by existing CSAT, pipeline-created, and QA-escape-rate dashboards.

Risks acknowledged. Morale during transition; short-term customer-facing quality dip; competitor poaching of departing staff. Mitigations: retention bonuses for pod leads; a public "AI-native" comms push in March.

Recommendation: approve Wave 1 at the March board meeting.

2. Risk-rate each assumption against known failure patterns

Now that the assumptions are surfaced, you want them scored against what's actually gone wrong in comparable rollouts. The Meta post-mortem is the freshest, most-documented example — the plan lost momentum as Meta found that its AI tools were not producing the gains executives had expected. The failure signal to watch for: internal figures showed a sharp increase in AI-assisted code, but a much smaller increase in product improvements reaching users. That gap — activity up, outcomes flat — is the single most common post-mortem finding, and step two prices it in.

Prompt
You now have a numbered assumption list from a workforce-automation plan. Score each assumption for risk using the rubric below, then return a ranked table (highest risk first).

Score each on 1-5 for:
- LIKELIHOOD OF BEING WRONG (5 = almost certainly wrong at scale)
- SEVERITY IF WRONG (5 = kills the plan / triggers customer or regulatory harm)
- REVERSIBILITY (5 = irreversible within 12 months, e.g., laid-off institutional knowledge)

Composite RISK = Likelihood × Severity × Reversibility (1-125).

Apply these known failure patterns from prior AI-workforce rollouts when scoring — flag any assumption that matches:
- "Activity-output gap": productivity metric (PRs, tickets closed, drafts produced) rises but downstream outcome (shipped features, resolved cases, revenue) does not. This is the Meta Project OT pattern — internal figures showed a sharp increase in AI-assisted work but a much smaller increase in results reaching users.
- "Extrapolation from a friendly domain": gains measured in one function (usually engineering) assumed to transfer to functions with different feedback loops (support, sales, QA).
- "Sunset-the-humans-too-soon": human-in-the-loop review layers timed to end before the error rate is actually characterized.
- "Consultant validation ≠ evidence": external partners endorsing the model is not the same as a pilot result.
- "Contingent second wave with soft gate": Wave 2 gated on Wave 1 "hitting KPIs" without pre-registered thresholds or independent measurement.
- "Institutional-knowledge cliff": the eliminated roles hold undocumented process knowledge that agents cannot recover.

Output columns: Rank | # | Assumption (short) | Likelihood | Severity | Reversibility | RISK | Failure pattern(s) matched | One-line "how you'd know early".

3. Turn the risk table into a lessons report and a go/no-go recommendation

Step three converts the analysis into something you can actually put in front of a board or a co-founder. Keep it short — the value is in the pre-mortem framing, not the length.

Prompt
Using the ranked risk table from the previous step and the original plan, write a 1-page "Pre-Mortem & Lessons Report" for the founder-CEO. Structure it exactly as follows, using markdown headers:

## Verdict
One of: PROCEED / PROCEED WITH CHANGES / PAUSE / KILL. Then one paragraph (max 4 sentences) justifying the call by referencing the top 3 risks by RISK score.

## The three assumptions most likely to sink this
Bullet each with: the assumption, why it's likely wrong (cite the failure pattern), and the earliest observable signal it's going wrong.

## Changes required before approval
A numbered list of 4-6 concrete edits to the plan. Each must be actionable this week (e.g., "replace the 60-day human-in-loop sunset with a rolling gate tied to a <2% escape-rate metric"). No vague "improve governance" bullets.

## Pilot design to de-risk Wave 1
Propose ONE 60-day pilot on the single function where the extrapolation is riskiest. Specify: which function, which agent scope, what metric proves it worked, what metric proves it failed, and the pre-registered go/no-go threshold.

## Lessons for the founder (keep after this project)
3 bullets. Durable operating principles, not project-specific. Written in the founder's own voice.

Keep the entire report under 600 words. No hedging language ("it may be worth considering"). Be direct.

Gotchas

  • Garbage-in problem. If the pasted plan is a slide-deck summary rather than the actual memo, the assumption extraction gets thin. Paste the longest version you have, including the appendix.
  • The model will over-anchor on Meta. Because the failure pattern list names Meta's Project OT explicitly, the risk-rater sometimes forces every plan into that mold. If your plan is for a 20-person company, not a 70,000-person one, tell the model that in step 2.
  • "Reversibility" is the score people fight about. Founders routinely mark laid-off roles as reversible ("we'll rehire"). They aren't — the institutional knowledge doesn't come back. Don't let the model soften this; if it scores reversibility below 3 for a headcount cut, push back in a follow-up turn.
  • Consultant-validated ≠ piloted. Meta ultimately cut about 10% of its workforce in May. The planned November phase did not go ahead. The gap between "two partners endorsed it" and "we ran a 60-day pilot" is the entire point of step 3's pilot design section — don't skip it to save time.
  • This is a decision aid, not a decision. The output is only as good as the plan you paste and the honesty of the assumptions inside it. Treat the report as a structured second opinion, not a verdict.

Time saved

A proper pre-mortem on a workforce-automation plan is usually a 2-week calendar item with 6-10 hours of the founder's own time — reading, arguing, drafting. This workflow gets you a defensible first draft in about 45 minutes. Call it 3-4 hours of founder time reclaimed per review, and — if the report catches one bad wave before you announce it — an amount of downstream cost that is not worth estimating on a public page. For context on why founders should treat this as a real risk, see the Reuters reporting on how the Meta plan actually unraveled.

What we got when we ran it

We executed every prompt above, in order, against the sample input (model: claude-sonnet-4-6). Outputs below are real and unedited apart from length truncation.

Step 1 output

| # | Assumption | Type | Evidence cited in the plan | Falsifiable? | |---|---|---|---|---| | 1 | The 3.1x increase in pull requests per engineer observed in the Q4 coding-assistant rollout can be extrapolated to Support, Marketing Ops, SDR, and QA functions at the same productivity multiple. | Productivity | "Our internal coding-assistant rollout in Q4 showed a 3.1x increase in pull requests per engineer." | Yes | | 2 | The Platform team can design, build, and deploy a production-ready stack of three agent frameworks plus a supervisor tool in 8 weeks. | Timeline | "a supervisor tool built by our Platform team (est. 8-week build)" | Yes | | 3 | The remaining 260-person workforce will successfully operate as talent-dense pods of 4-6 humans supervising fleets of AI agents without requiring significant retraining, role redesign, or additional management infrastructure. | Org-design | "remaining team operating as 'talent-dense pods' of 4-6 humans supervising fleets of AI agents" | Yes | | 4 | Post-Wave-1 output will reach at least 85% of pre-reorg levels by end of Q3, sufficient to trigger Wave 2 on schedule in October. | Productivity | "contingent on Wave 1 hitting >85% of pre-reorg output KPIs by end of Q3" | Yes | | 5 | Existing dashboards (CSAT, pipeline-created, QA-escape-rate) are adequate and sufficient to measure whether AI agents are maintaining organizational output quality at the required threshold. | Quality | "Quality will be measured by existing CSAT, pipeline-created, and QA-escape-rate dashboards" |

…output truncated at 250 words.

Step 2 output

| Rank | # | Assumption (short) | Likelihood | Severity | Reversibility | RISK | Failure pattern(s) matched | How you'd know early | |------|---|---|---|---|---|---|---|---|---| | 1 | 1 | 3.1x coding gain transfers to Support / SDR / QA / Mktg Ops | 5 | 5 | 5 | 125 | Activity-output gap; Extrapolation from a friendly domain | Track resolved cases / pipeline-accepted / escaped defects — not tickets closed or drafts produced — in week 1 of Wave 1 | | 2 | 11 | AI agents perform Tier 1, SDR, Mktg Ops, Junior QA at acceptable quality | 5 | 5 | 4 | 100 | Activity-output gap; Extrapolation from a friendly domain; Institutional-knowledge cliff | CSAT and QA-escape-rate diverge from ticket-volume and PR-count within first 30 days post-Wave-1 | | 3 | 6 | Human-in-the-loop can safely sunset at 60 days | 5 | 4 | 5 | 100 | Sunset-the-humans-too-soon | Error-rate trend at day 45 is still declining rather than having plateaued for ≥2 weeks | | 4 | 10 | Departing staff won't materially harm competitive position | 4 | 5 | 5 | 100 | Institutional-knowledge cliff | Monitor competitor job postings and customer-relationship assignments of departing AEs / senior support staff within 60 days of Wave 1 | | 5 | 14 | Productivity gains are durable as task complexity rises | 5 | 4 | 4 | 80 | Activity-output gap; Extrapolation from a friendly domain | Agent

…output truncated at 250 words.

Step 3 output

Verdict **PAUSE** The plan's financial and headcount logic rests entirely on one extrapolation — that a 3.1x coding productivity gain transfers to Support, SDR, QA, and Marketing Ops — and that claim scores 125/125 on the risk rubric. No pilot data exists for any of the four target functions. Compounding this, the 60-day human-in-the-loop sunset permanently destroys the institutional knowledge needed to catch agent errors before they reach customers or kill pipeline. Approving Wave 1 in March means the board is betting $14.2M in projected savings — and 160 people's jobs — on an analogy, not evidence. --- ## The three assumptions most likely to sink this - **3.1x gain transfers across functions.** It won't. Pull requests are a self-contained, fast-feedback loop; a resolved support case, an accepted sales meeting, and a zero-escape QA build are not. This is textbook *extrapolation from a friendly domain* — the Meta Project OT pattern, where internal activity metrics looked strong while user-facing outcomes lagged badly. **Early signal:** CSAT and QA-escape-rate diverge from ticket-volume within the first 30 days of any pilot. - **AI agents perform Tier 1, SDR, and QA at acceptable quality.** These functions require contextual judgment, relationship memory, and error-recovery that current agent frameworks handle poorly at scale. The plan cites no benchmark, no vendor SLA, and no shadow-run result. **Early signal:** human-override frequency on the supervisor tool exceeds 15% in week two — meaning humans are already doing the job agents were hired for. - **Human-in-the-loop sunsets safely at

…output truncated at 250 words.

Source: news.google.com

More for Founder & Operator professionals →

Get the next one in your inbox

One daily brief. Every story gets a hype verdict.

No spam. Unsubscribe anytime.

Exact prompts included · Untested steps are marked · Corrections are public