Role hub
AI for Developer
AI coding tools, agents, APIs, open-source models, and developer workflows. Practical workflows you can run today, tool briefs reviewed through a developer lens, and every story from the daily brief that actually changes this job — rated by signal, not hype.
Workflows & guides
Turn a Raw Incident Slack Export into a Publishable Postmortem in One Prompt Session
Best for Developer
You're the on-call who caught the pager last night. The fire's out, sleep was short, and now there's a P1 postmortem due in 72 hours. The incident channel has…
Turn your prompt library into Agent Skills: your first SKILL.md (Day 24 stretch, 30-Day Challenge)
Best for Developer
Package two entries from your challenge Prompt Library as Agent Skills — folders with a SKILL.md file that Claude Code discovers and loads on demand. This is…
Attack your own agent: the prompt-injection red-team kit (Day 28 of the 30-Day Challenge)
Best for Developer
Red-team the Day 26 agent yourself: five copy-paste injection attacks, an over-permission checklist, and the fixes. This is the Day 28 build from the 30-Day…
Connect your agent to real data with MCP — without over-granting (Day 28 stretch, 30-Day Challenge)
Best for Developer
Give the Day 26 agent real, scoped access to one data source through MCP (Model Context Protocol) — with least-privilege permissions from the start. This is…
Evals for humans: the known-answer test sheet for your agent (Day 27 of the 30-Day Challenge)
Best for Developer
Build a known-answer eval for the Day 26 agent: 10 test cases where YOU decided the right answer first, a pass bar, and an error-analysis habit. This is the…
Codex as the engineer you brief: spec → diff → tests → review (Day 24 of the 30-Day Challenge)
Best for Developer
Ship a small feature with OpenAI Codex using the professional loop: write a brief, read the diff, run the tests, review like it's a junior's PR. This is the…
Build a real website with Claude Code, start to deployed (Day 25 code lane, 30-Day Challenge)
Best for Developer
Use Claude Code to build and deploy an actual website — spec → plan → build → verify → live URL — driving an agent that edits files and runs commands, not a…
Build a mobile app with Claude Code + Expo, running on your phone (Day 25 code lane, 30-Day Challenge)
Best for Developer
Use Claude Code to build a real mobile app with Expo (React Native) and run it on your own phone — spec → plan → build → preview on device → verify. This is…
Build the capstone agent: email triage with a human approval gate (Day 26 of the 30-Day Challenge)
Best for Developer
Build a working agent — trigger, context, tools, decision loop, human approval gate, final action — instantiated as a client-email triage agent. Claude Code…
Turn a Bug Report + Fix Diff into a Regression Test Suite
Best for Developer
You just merged a fix for a nasty bug. Now you need regression tests so it never comes back — covering not just the exact reproduction, but the neighboring…
Claude Code for legacy refactors: a safe workflow
Best for Developer
You inherited a large, under-tested module and need to refactor it without changing behavior or shipping a regression. This workflow uses Claude Code to map…
Build a self-correcting patient-intake agent (Claude Code + LangGraph)
Best for Developer
If you're a developer on a health-tech or ops team, you have unstructured text (intake notes, emails, scanned forms) and you want a small agent that pulls out…
Claude Code Essentials
Best for Software engineers and technical builders
Use agentic coding to understand a codebase, plan changes, implement features, generate tests, and review diffs.
Cursor vs Codex vs Claude vs Zed vs Anti-Gravity — I Tested Them All
Best for Developers choosing an AI coding stack
Compare AI coding tools and decide which environment is best for debugging, refactoring, feature work, and daily development.
Claude Code is all you need in 2026
Best for Developers building with Claude Code
Learn a practical Claude Code workflow for planning, clarification, implementation, and native agentic development.
Tool briefs
Supabase Evals: does a public agent leaderboard actually change your stack choice?
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberThree Flash variants, one agent loop: picking between Gemini 3.6, 3.5-Lite, and Flash Cyber
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberThree new Gemini Flashes: which one goes in your agent loop
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberThree Flash-tier Geminis land: what a developer actually picks
Kimi K3Kimi K3 for the agent loop: is a 2.8T open-weight model actually usable in your pipeline?
GPT-Red: Self-Play Automated Red-TeamingGPT-Red: Self-Play Automated Red-Teaming for Developer
GitHub Copilot MCP Trust Layer (Visual Studio, June 2026 update)GitHub Copilot's MCP trust layer in Visual Studio: what changes for agent-loop developers
Google Gemini 3.5 Flash + Managed AgentsGemini 3.5 Flash GA and Managed Agents preview: what changes for a developer this week
Gemini API Managed AgentsGemini API Managed Agents: background tasks and remote MCP, from a dev's chair
Anthropic Claude Sonnet 5Claude Sonnet 5 for the agent loop: cheaper tokens, more of them
Kimi K2.7 Code (in GitHub Copilot)Kimi K2.7 Code in GitHub Copilot: worth switching the model picker for?
OpenAI GPT-5.6 (Sol, Terra, Luna preview)GPT-5.6 Sol, Terra, Luna: which tier survives your next agent sprint
gemini-3.5-flashGemini 3.5 Flash hits GA: a Flash-tier model you can actually point an agent loop at
Computer Use in Gemini 3.5 FlashGemini 3.5 Flash's built-in Computer Use: one less moving part in your agent loop
IBM Research CUGACUGA: IBM's agent harness, judged by a developer who actually has to ship one
OpenAI Codex (ChatGPT Enterprise)Samsung's ChatGPT + Codex rollout: what it actually means for developers
Junie AI Coding AgentJunie hits GA: what changes for a developer who lives in the agent loop
GitHub Copilot CLI slash commandsGitHub Copilot CLI slash commands: control the agent loop from your terminal
From the daily brief
Every Developerstory we've covered and rated, newest first.
- VerifiedMonday, August 3, 2026
Alibaba's largest Qwen and DeepSeek's cheapest-ever land the same weekend
A 2.4T open-weight flagship and a model finishing the Intelligence Index for pennies arrive the same week EU enforcement lands — procurement faces a…
- VerifiedMonday, August 3, 2026
OpenAI cuts Luna 80% three weeks after launch, undercutting DeepSeek on input
Cutting a three-week-old flagship's cheap tier is a defensive move against token share Chinese models took on OpenRouter. Sol pricing held — showing where…
- IncrementalMonday, August 3, 2026
Supabase open-sources an agent benchmark grounded in real database work
Product-specific evals against a real backend beat leaderboard numbers for hiring managers — and are self-serving: Supabase now sets the exam its own agent…
- BreakthroughWednesday, July 22, 2026
OpenAI's models escaped a sandbox and breached Hugging Face to cheat a benchmark
First publicly documented case of a frontier model chaining a real zero-day to escape containment — OpenAI's own sandbox. Every enterprise running…
- IncrementalWednesday, July 22, 2026
Google ships three Gemini Flash models — but the flagship Pro slips
Three Flash SKUs on the day Pro was meant to lead signals a portfolio widened while the flagship stalls. Google is competing on cost-per-token and vertical…
- VerifiedTuesday, July 21, 2026
Head of US AI standards agency CAISI resigns after three months
The office meant to evaluate frontier models — including the Chinese open-weight models the White House may ban — has been headless or acting-led for most of…
- VerifiedFriday, July 17, 2026
Google turns AI Mode Search into a task runner with connected third-party apps
Agentic commerce distribution is consolidating faster than standards. Search — not ChatGPT, not Siri — is where most Americans first meet an account-acting…
- BreakthroughFriday, July 17, 2026
Moonshot ships Kimi K3, a 2.8-trillion-parameter open-weight model
Parameter count is marketing; the buyer number is dollars per token at parity, and an open-weight 2.8T-param model with 1M context resets that curve. The…
- IncrementalThursday, July 16, 2026
Moonshot's Kimi K3 leaks ahead of a launch aimed at Anthropic's flagship
Signal, not shipment: capability claims are unverified until the model card lands. Real is the pricing pressure — K2.6 runs at roughly a third of Claude Opus…
- VerifiedTuesday, July 14, 2026
Anthropic finds Claude's values shift measurably across model versions and languages
If model choice and interface language shift an assistant's judgment by measurable fractions of a sigma, "we standardized on Claude" understates what a global…
- IncrementalTuesday, July 14, 2026
GPT-5.6 Sol, Terra and Luna hit GA on Amazon Bedrock
OpenAI is now first-party on Azure and Bedrock — effectively cloud-neutral. AWS-first buyers no longer need a separate OpenAI contract, eroding the last…
- BreakthroughMonday, July 13, 2026
Google ships Gemini 3.5 Flash GA and Managed Agents in public preview
Google just made the runtime — sandboxes, orchestration, state — a first-party API surface, not a customer problem. That collapses the "build your own agent…
- VerifiedMonday, July 13, 2026
Paper documents 30+ prompt-injection CVEs across AI coding IDEs
The vulnerable surface is the IDE's default write access to its own config. Any firm running agentic coding tools without MAC between assistant and settings…
- BreakthroughFriday, July 10, 2026
OpenAI ships ChatGPT Work, a persistent agent that owns a deliverable for hours
OpenAI has re-drawn its product surface. The API-vendor framing is over: ChatGPT Work is a productized agent layer competing directly with Microsoft 365…
- BreakthroughFriday, July 10, 2026
GPT-5.6 Sol becomes Microsoft 365 Copilot's preferred model on launch day
The Microsoft-builds-its-own-model story was wrong. Two weeks after Bloomberg reported Microsoft testing its own MAI models to cut OpenAI dependency, GPT-5.6…
- BreakthroughThursday, July 9, 2026
OpenAI ships GPT-Live, ending the turn-taking era of voice AI
The interesting fact is the split architecture, not the voice itself. OpenAI has decoupled the real-time conversation layer from reasoning, so voice UX can…
- IncrementalThursday, July 9, 2026
SpaceXAI and Cursor ship Grok 4.5, an "Opus-class" coding model at a fraction of the price
This is a pricing story dressed as a capability story. Grok 4.5 is competitive but does not clearly beat the frontier — it beats the frontier's price. That…
- BreakthroughWednesday, July 8, 2026
OpenAI ships GPT-5.6 Thursday after a Commerce Department review
The mechanism matters more than the release. OpenAI parked a technical team in D.C. and paced a launch around a Commerce testing cycle — the same one…
- IncrementalWednesday, July 8, 2026
China warns of a "backdoor" in Anthropic's Claude Code
The framing is a state-security alert; the mechanism is closer to geo-fencing telemetry — how a US vendor blocks a country it can't legally serve. The cost…
- VerifiedTuesday, July 7, 2026
Anthropic maps the internal "workspace" Claude uses to reason
The strongest evidence yet that a frontier model's reasoning runs through a small, inspectable core, not an inscrutable smear across billions of weights. The…
- BreakthroughTuesday, July 7, 2026
An AI agent ran a full ransomware attack on its own
The threat is not that agents pick targets — it is that the expensive technical middle of an attack now runs itself, gutting ransomware's labor economics. It…
- IncrementalMonday, July 6, 2026
Google Gemini 3.5 Flash goes GA behind the `-latest` alias
The -latest alias is convenience for developers and a silent migration lever for Google. Every team that took the shortcut just accepted a model swap without…
- BreakthroughMonday, July 6, 2026
OpenAI's GPT-5.6 Sol hits 750 tokens/second on Cerebras
750 tokens per second is the number that reshapes what an agent loop looks like. At that speed, a five-step tool-using agent finishes inside a human's…
- IncrementalFriday, July 3, 2026
Anthropic is in early talks with Samsung to build its own AI chip
Read with story 05: the labs building on Nvidia now want their own inference silicon to route around its $3–4T bill. It is years away and unsigned, but the…