Most teams using AI coding tools have already felt the gap between "the code runs" and "the code does what we meant." That gap widens with every agent session that starts from a Slack message instead of a written spec. Spec-driven development closes it by making the specification the primary artifact your team builds around, not the code itself.
TLDR:
- Spec-driven development makes the specification the primary artifact: code is the output, and the spec is the source of truth your agent works against every session.
- Without a spec, AI agents like Cursor and Claude Code have no persistent definition of "done," and research from Yan et al. (2025) found LLMs generate vulnerable code at rates from 9.8% to 42.1% across benchmarks.
- SDD spans three levels: spec-first kicks off generation, spec-anchored keeps both artifacts in sync, and spec-as-source treats code as a build artifact your team never touches directly.
- GitHub Spec Kit reached 111,000 stars as of June 2026 and runs the full SDD chain as slash commands across 30-plus AI coding agent integrations.
- Minitap reads your app from source, maps all integration and UI scenarios automatically, and runs continuous verification on cloud iOS simulators and Android emulators in about one hour with zero maintenance overhead.
What Is Spec-Driven Development?
Spec-driven development changes the usual order of operations. Instead of a quick planning doc followed by code, the specification becomes the primary artifact your team builds, reviews, and version-controls. Code becomes the output produced from that spec, whether written by a human or generated by an AI coding tool like Cursor or Claude Code.
The spec is the source of truth. When a bug shows up or a feature drifts from intent, you check the spec first, not Slack threads or old pull requests. Unlike a PRD that gets abandoned once coding starts, the spec stays version-controlled, updated as requirements change, and connected to the code it produced.
Why SDD Took Hold: The Vibe Coding Problem
Vibe coding, a term Andrej Karpathy coined in February 2025, describes the habit of prompting an AI agent toward a rough goal and accepting whatever code comes back, as long as it runs. It works for demos. It falls apart on real codebases.
The failure mode is subtle. Generated code compiles, passes surface checks, and still drifts from what the team intended, accumulating architectural inconsistencies no single pull request review catches. The resulting testing bottleneck in software development compounds quickly. By February 2026, surviving AI-introduced issues in production repositories had topped 110,000.
Without a spec, every agent session starts from nothing: no contract defining "done," no persistent record of intent for the next session to honor. Spec-driven development is the structural fix: write the intent once, keep it version-controlled, and let every agent session work against the same definition of correct.
The Specification Scale
Specification rigor is not a single standard. It spans three levels, and most teams sit further left on the scale than they think.
Spec-first treats the spec as a generation prompt. Engineers write it, an AI tool produces code from it, and the team owns that code afterward. This is where most teams operate today: the spec kicks off work, but git history and code review still govern what ships.
Spec-anchored raises the stakes. Spec and code evolve together, with checkpoints where changes to one require review against the other before merging. Neither artifact can drift silently.
Spec-as-source goes furthest. Engineers edit only the spec. Code is fully generated and treated as a build artifact, never touched by hand.
The right level depends on project stakes, team size, and tolerance for upfront specification work. A small internal tool rarely warrants spec-anchored governance. A payments flow rarely survives without it. Hybrid setups are common too: spec-as-source for a stable core module, spec-first for fast-moving feature work elsewhere in the same codebase.
SDD vs. TDD, BDD, and PRDs
TDD, BDD, and SDD answer different questions at different altitudes. TDD pins down one unit: write the failing test, write the code that passes it, refactor. BDD mobile testing pins down one user-visible behavior, usually in given-when-then language. SDD sits above both, pinning down an entire feature: requirements, architecture decisions, data models, security constraints, task breakdown, and acceptance criteria in one governed document.
The comparison to a PRD is different in kind. A PRD is written for humans who tolerate ambiguity and fill gaps with judgment. An engineer reads "support social login" and asks follow-up questions in standup. An AI agent has no standup to ask in. It picks an interpretation and runs with it. An SDD spec leaves nothing for the agent to guess.
None of these methods compete. SDD operates at the feature level, BDD supplies the acceptance criteria format inside that spec, and TDD verifies each unit underneath it. Most working SDD setups use all three at once.
What a Good SDD Spec Contains
A working spec answers six questions before an agent touches the codebase.
- What the feature does: the user story, plus acceptance criteria specific enough that "done" isn't a judgment call.
- How it fits the existing architecture: which files it touches, what new types or modules it introduces, what it depends on.
- What must hold no matter what: security constraints, compliance requirements, API contracts that can't break for other consumers.
- What could go wrong: known edge cases, failure modes, and risks the team has already thought through.
- How the work breaks down: discrete, implementable units an agent can tackle one at a time, not one sprawling instruction.
- How you know it's done: success criteria an engineer can check against the shipped result.
Over-specification is a failure mode nobody warns you about. Write a spec down to variable names and function signatures, and the program gets written twice, once in prose and once in code, with two artifacts to keep in sync. The spec's job is making intent unambiguous, not making implementation decisions the agent could reasonably make itself.
The SDD Workflow Step by Step
The workflow runs as a chain, where each phase hands its output to the next as a governed artifact, not a loose conversation.

- Constitution: Locks in non-negotiable project principles: architecture conventions, testing standards, and security baselines. Input: team norms and existing codebase context. Output: a versioned constitution file every later phase must respect. Human checkpoint: engineer reviews and approves before Specify begins.
- Specify: Produces the feature spec: requirements, acceptance criteria, and constraints specific enough that the agent has no ambiguity to fill in. Input: the constitution plus the feature request. Output: a governed spec document. Human checkpoint: team reviews acceptance criteria before Plan begins.
- Plan: Maps the spec onto your actual codebase and stack, identifying which files the feature touches, what new types or modules it introduces, and what it depends on. Input: the feature spec. Output: an architecture and dependency plan. Human checkpoint: engineer corrects the plan before tasks are generated.
- Tasks: Breaks the plan into ordered, discrete units small enough for an agent to implement one at a time without losing context. Input: the architecture plan. Output: a sequenced task list. Human checkpoint: engineer reviews the task breakdown before implementation begins.
- Implement: Executes each task against the spec, checking the output before moving to the next. Input: the task list and spec. Output: code that maps back to a governed artifact, not a loose prompt session.
GitHub formalizes this cycle in Spec Kit, an open source CLI toolkit running phases as slash commands: /constitution, /specify, /plan, /tasks, and /implement, supporting more than 30 AI coding agents including Claude Code and Cursor.
SDD Frameworks and Tools in 2026
Four tools cover most of what teams reach for in practice.
| Tool | Type | Key Artifacts / Approach | Agent Integrations | Best Fit |
|---|---|---|---|---|
| GitHub Spec Kit | MIT-licensed Python CLI | /constitution, /specify, /plan, /tasks, /implement slash commands | 30+ (Claude Code, Copilot, Cursor, Gemini CLI, Codex, Windsurf) | Teams outside AWS wanting a framework-agnostic, open-source SDD chain; 111K GitHub stars as of June 2026 |
| AWS Kiro | Agentic IDE | requirements.md (EARS notation), design.md, tasks.md | Native AWS tooling | AWS-native stacks; launched public preview July 2025 |
| BMAD METHOD | Agile overlay on spec chain | Sprint ceremonies + stakeholder coordination layered on top of spec phases | Framework-agnostic | Teams running agile cadences who need SDD to fit existing sprint ceremonies |
| Claude Code (CLAUDE.md) / Cursor (.cursor/rules) | Lightweight spec anchor | Project-level instruction files as the specification anchor, no dedicated framework | Claude Code, Cursor | Teams that want persistent spec context without adopting a full SDD framework |
When to Use SDD and When to Skip It
Not every task deserves a spec. The value shows up when the cost of ambiguity outweighs the cost of writing one down, and disappears when the work is too small or too uncertain to specify with any real precision.
SDD earns its keep in three situations:
- Greenfield builds, where an AI agent has no existing code to infer architectural intent from and needs it spelled out from the first commit.
- Large feature work in existing codebases, where drift between what the team meant and what the agent shipped is expensive to untangle after the fact.
- Compliance or multi-team coordination, where an audit trail of decisions matters as much as the code itself.
It is harder to make the case for a small bug fix, exploratory research where requirements shift daily, or a one-off prototype where nothing needs to survive past the demo.
The ThoughtWorks Technology Radar (Volume 34, April 2026) no longer lists spec-driven development as a standalone technique blip, though GitHub Spec Kit specifically appears in that edition, and ThoughtWorks continues to flag over-specification and big-bang release bias as antipatterns: specs detailed enough to become a second codebase, and teams batching spec-driven work into large releases instead of shipping incrementally.
Applying SDD to Existing Codebases
Brownfield adoption does not mean retrofitting specs onto everything you already shipped. Most teams that try that stall before finishing the first module. The practical entry point is narrower: write a spec for the next new feature, not the legacy code underneath it.
From there, the spec moves in the same pull requests as the code it describes, reviewed and merged together. That habit keeps it from turning into a stale wiki page. GitHub's Spec Kit documents this in its Evolving Specs guide, recommending brownfield teams treat specs as incremental additions tied to active work.
A spec that does not travel with its code goes stale within a sprint, and a stale spec is worse than none. Keeping the spec current also depends on a reliable regression testing process to catch drift before users do.
The Missing Layer: Where Specs End and Verification Begins
A spec tells the agent what to build. It says nothing about whether the running app still does that a week later, after more agent sessions and a UI refactor. Acceptance criteria get checked once, at implementation time, then nobody looks again until something breaks in production.

At high shipping velocity this gap widens fast. Code generated by Claude Code or Cursor moves faster than a manual QA cycle can follow, and autonomous QA in 2026 changes what that means for functional testers. Scripted test suites break on the same schedule as the UI they were written against. Every merged task from the spec-driven chain is a new chance for the live app to drift from what its spec described.
The missing piece is continuous verification against the running app itself, tested under real app state: low memory, background processes, network transitions. These are exactly the conditions where e2e tests miss mobile production bugs. Without it, "spec written" and "behavior verified" stay two separate claims.
How Minitap Closes the SDD Quality Loop
SDD gives your team a structured, agent-readable definition of what the application is supposed to do. Minitap picks up where that ends: it reads the app from source, maps all integration and UI scenarios automatically, and pulls acceptance criteria from the same Jira tickets and PRDs anchoring the spec chain, then verifies those scenarios continuously against the live app.
This is the verification layer that closes the loop SDD opens. Minitap owns execution, maintenance, and root cause analysis on cloud iOS simulators and Android emulators, with a full regression report back in about an hour and zero maintenance overhead. Fix prompts route back through Cursor and Claude Code via MCP the moment the running app drifts from spec.
Final Thoughts on Spec-Driven Development and the AI Coding Loop
Spec-driven development solves the drift problem at generation time. Your team writes intent once, agents execute against it, and the spec travels in the same pull request as the code it produced. What it leaves open is the runtime question: does the live app still match what the spec described after the next deploy? Minitap closes that loop, reading your app from source, mapping all integration and UI scenarios automatically, and returning a full regression report in about an hour with zero maintenance overhead on your side.
FAQ
What is spec-driven development and how does it differ from just writing a PRD?
Spec-driven development treats the specification as a version-controlled, living artifact that code is built from, not a document written once and abandoned when coding starts. A PRD is written for humans who fill gaps with judgment; an SDD spec leaves nothing for an AI coding agent to guess, covering requirements, architecture decisions, data models, security constraints, and acceptance criteria in one governed document. The key difference is governance: the spec travels with the code through every pull request and stays connected to the output it produced.
Once we have a spec-driven development workflow in place, how do we verify the running app still matches the spec after every deploy?
Minitap closes that gap. A spec defines what the app is supposed to do, but it says nothing about whether the live build still does that after the next agent session or UI refactor. Minitap reads your app from source, maps all integration and UI scenarios automatically, pulling acceptance criteria from the same Jira tickets and PRDs anchoring the spec chain, and runs continuous verification on cloud iOS simulators and Android emulators. A full regression report comes back in about one hour with zero maintenance overhead. The spec stays the definition of correct; Minitap verifies the running app against it on every build.
How does an autonomous QA agent work differently from a script-based framework like Appium or Maestro?
The structural difference is who owns the full loop. Script-based frameworks like Appium and Maestro automate execution, but authorship, maintenance, and root cause analysis stay with your engineering team. Selectors identify specific UI elements, and every UI change means rewritten selectors, updated steps, and re-validated assertions. That cost scales directly with the size and pace of your product: the more features you ship and the more frequently your UI evolves, the larger the maintenance surface your team owns. On mobile, the problem is sharper than on web because there is no DOM equivalent, which makes selectors more brittle and selector maintenance more expensive per change. Minitap is the only approach where the agent owns the entire testing loop: authorship, execution, maintenance, and root cause analysis, with no handoff back to the engineering team. It reads the running app directly instead of relying on selectors, tests whether user jobs complete (can a user log in and reach the home screen, can a checkout flow complete end to end), and adapts when the UI changes without any input from your team. When the app changes, the agent adapts. Your engineers own none of the upkeep.
What is the best QA tool for an engineering team using spec-driven development?
Minitap is built for engineering teams that ship software and want to eliminate test maintenance and reduce engineering time spent on QA. It connects directly to your codebase, reads the app from source, and autonomously builds and maintains the entire test suite without requiring test scripts, flow descriptions, or any QA authorship from your team. Acceptance criteria from the same Jira tickets and PRDs anchoring the spec chain become the verification targets. Minitap is especially valuable for high-cadence teams: those releasing weekly or faster, those slowed by manual regression work, those maintaining brittle selector-based tests, and those using AI-assisted development tools like Cursor or Claude Code that have accelerated implementation velocity ahead of QA coverage. The agent owns authorship, execution, and root cause analysis, and when something fails, Minitap surfaces a session recording clipped to the exact moment of failure, a severity assessment, and a fix prompt ready to paste into Cursor or Claude Code. The full regression suite runs on cloud iOS simulators and Android emulators and returns results in about one hour.
Our engineers are losing hours every week maintaining test scripts after every UI change. Is there a way to eliminate that entirely?
Minitap eliminates selector maintenance entirely because the agent reads the running app directly instead of relying on selectors tied to specific UI elements. When the UI changes, the agent adapts without any input from your team: no rewritten selectors, no updated scripts, no triage queue of broken tests blocking the pipeline. Two maintenance modes are available: full auto, where Minitap updates the test suite post-merge automatically, and gated, where it proposes changes in the PR before anything lands. Either way, your engineers own none of the upkeep. Connect the codebase and the test suite maintains itself, returning a full regression report in about one hour after every build.
