When a team first hears about Spec-Driven Development (SDD), building software from a written specification that the AI works from, the first question is usually whether it is yet another methodology for the list. They already had TDD, they had tried BDD, and now a third acronym arrives with "Driven" in the middle.
The question makes sense, but it starts from a wrong premise: that all three compete for the same place. They don't. Each one answers a different question, and when an agent writes the code, all three questions are still there. What changes is who reads the answer.
The thesis of this piece: SDD does not replace TDD or BDD; it needs them. BDD brings the way criteria are written, with concrete examples. TDD brings the cycle that proves they are met, with a test that fails before the code exists. SDD adds that this specification is what the agent reads and the source of truth for the work. Remove any of the three and you leave a gap the agent will not fill on its own.
Three different questions
Before comparing them, it helps to separate which question each one answers. They are easy to confuse because all three talk about specifying before programming.
| Practice | Question it answers | What it produces | What it adds when an agent writes |
|---|---|---|---|
| BDD (Behaviour-Driven Development) | How do we describe what the system should do? | Concrete examples, readable by people and by machines | Unambiguous criteria that can become scenarios |
| TDD (Test-Driven Development) | How do we prove the code meets it? | A test that fails first and passes afterwards | A check the agent does not control |
| SDD (Spec-Driven Development) | What does the agent read before writing? | A specification that is the input to the work and the source of truth | The context and scope of the task, written down and reviewable |
onext's own analysis, based on Cucumber, the spec-kit specification and Birgitta Böckeler's analysis (accessed on 10 October 2026)
BDD: how the criterion is written
Cucumber, the tool built to support BDD, organises it into three practices. In discovery, structured conversations around real-world examples of the system from the users' perspective. In formulation, each example is written as structured documentation, in "a medium that can be read by both humans and computers". In automation, that executable specification guides the implementation.
That way of writing has not disappeared with agents: it has come back. Birgitta Böckeler, of Thoughtworks, analysed three SDD tools and describes how Kiro structures requirements as user stories with acceptance criteria in GIVEN… WHEN… THEN… format, the one BDD made popular. A criterion written like that can be read in a meeting and also run as a test.
The same analysis carries the opposite warning: in one of her trials, the requirements document turned a small bug fix into four user stories with sixteen acceptance criteria. The format helps you be precise; it does not decide how precise you need to be. That decision is still a person's, and we argued it in more detail in MVP versus specification.
TDD: the proof that it is met
This is the part most often lost when people talk about SDD. The spec-kit specification, GitHub's kit for working with SDD, calls it the "Test-First Imperative" and writes it as a rule: all implementation must follow strict TDD, and no implementation code is written until the tests are written, "validated and approved by the user" and confirmed to fail. Elsewhere it sums it up in one sentence: acceptance scenarios become tests.
Kent Beck, who popularised TDD, describes in "Augmented Coding: Beyond the Vibes" how he tried to get his agent to work that way. His instructions were explicit: always follow the red, green, refactor cycle; write the simplest failing test first; one test at a time. And he lists three signs that the agent was going off track: loops, functionality he hadn't asked for, and any indication that it was cheating, for example by disabling or deleting tests.
Risk An agent that can edit the tests can make them "pass" without meeting anything. Changes to tests are changes to the criterion the rest of the code is judged against: they deserve a separate review, and before the review of the code.
There is a nuance worth not skipping: a test existing does not mean it tests what it should. A test written by looking at the code that is already there describes what that code does, not what it should do; we explain it in AI-generated tests describe your code, they don't test it. That is why order matters: the test comes from the criterion and fails before the code exists.
SDD: what changes when the reader is an agent
What SDD adds to the other two is not a new testing technique. It is a change of reader. In TDD and BDD, the one who reads the criterion is a person who is about to program. In SDD, it is an agent, and the specification becomes the input to the work: the context, the scope and what counts as done.
Böckeler distinguishes three levels. In the first, the specification is written beforehand and used for the task at hand (spec-first). In the second, it is kept afterwards to keep evolving that feature (spec-anchored). In the third, the specification is the main source file and the person no longer touches the code (spec-as-source). For a team starting out, the first level is the natural starting point, and that is where combining it with TDD pays off most.
And this is the underlying reason for the thesis. Böckeler puts it in her own words: even with all those files, templates, prompts, workflows and checklists, she frequently saw the agent ultimately not follow all the instructions. She also saw it go overboard with some. A specification the agent reads but nothing checks depends on the agent following it. A test that fails for as long as the behaviour does not exist does not.
How they fit into a task
Put in order, the three practices do not overlap: each one covers a stretch. This sequence is a proposed criterion, not a rule from any of the sources:
- Examples before requirements. Business and development agree on two or three concrete cases of what should happen, including one that should not (BDD, discovery).
- Criteria in the specification. Each example is written as a verifiable acceptance criterion, in GIVEN/WHEN/THEN or an equivalent format, inside the specification the agent will read (BDD, formulation; SDD).
- Failing tests, approved by a person. Each criterion produces a test; someone checks that it tests what the criterion says and that it fails (TDD, red).
- The agent implements. With the specification as context and the tests as the target, without permission to change them unreviewed (SDD and TDD, green).
- Review what the tests do not cover. What has gone green is not reviewed again line by line; what gets reviewed is whatever no criterion covered and any change to the tests.
The third step is what turns the specification into a gate rather than a document: we develop it in your spec is already an eval. And the point where a person signs off that specification, before the agent generates anything, in the specification is where you sign off.
The disagreement: is it closer to TDD or to MDD?
It would be more comfortable to present SDD as the natural evolution of TDD and BDD. Böckeler doesn't quite see it that way. She acknowledges that many people draw analogies between SDD and TDD or BDD, but suggests looking at another parallel, particularly for the spec-as-source level: Model-Driven Development (MDD), where the models were the specifications and a generator produced the code. MDD never took off for business applications, and she wonders whether that level of SDD might end up with the downsides of both MDD and language models.
Her conclusion is even more uncomfortable: she wonders whether some tools are feeding agents our existing workflows too literally, amplifying challenges that already existed, such as review overload. She uses a German word for it, Verschlimmbesserung: making something worse in the attempt to make it better.
There is no need to settle that debate here, but there is a practical consequence to draw. The risk she describes grows the further the specification drifts from something that can be checked: many files, a lot of prose, no failing test. Combining it with TDD is exactly what keeps SDD on the verifiable side. It is the same argument we started from when writing about Spec-Driven Development and controlled code.
And if that specification lives in a repository with an assistant such as Claude Code, the other half of the question is which rules are suggested to the agent and which are enforced; we cover it in CLAUDE.md is context, not control.
Frequently asked questions
What is the difference between SDD, TDD and BDD?
They answer different questions. BDD (Behaviour-Driven Development) is about how the expected behaviour is described: with concrete examples, agreed between business and development and written so that both people and machines can read them. TDD (Test-Driven Development) is about how you prove the code meets it: you first write a failing test and then the minimum code that makes it pass. SDD (Spec-Driven Development) is about what the agent that writes the code reads: the specification is the input to the work and the source of truth.
Does SDD replace TDD?
No. The spec-kit specification itself, GitHub's kit for SDD, includes TDD as a rule: no implementation code before the tests have been written, approved by the user and confirmed to fail. Without a test that fails first, the specification guides the agent, but nothing proves it was met.
Do specifications have to be written in GIVEN/WHEN/THEN format?
It isn't mandatory, but it helps. Birgitta Böckeler describes how Kiro structures requirements as user stories with acceptance criteria in GIVEN… WHEN… THEN… format. That is the format BDD made popular, and its advantage is that every criterion can become an executable scenario. The risk runs the other way: Böckeler saw a small bug fix turned into four user stories with sixteen acceptance criteria.
Why can an agent skip the specification?
Because reading it does not guarantee following it. Böckeler reports that, even with templates, workflows and checklists, she frequently saw the agent not follow all the instructions, and she also saw it go overboard by following one too eagerly. That is why you need something the agent does not control: a test that fails for as long as the behaviour does not exist.
What do I do if the agent changes or deletes tests to make them pass?
Treat it as a warning sign, not a detail. Kent Beck lists it among the three signs that the agent is going off track: any indication that it is cheating, for example by disabling or deleting tests. In practice, changes to tests should be reviewed separately from the code and more carefully, because they are the criterion everything else is judged against.
Is SDD the same as Model-Driven Development?
No, but some see the resemblance. Böckeler points out that, especially at the level where the specification is the only source a person edits (spec-as-source), another important parallel is MDD: the models were specifications from which code was generated. MDD never took off for business applications, and she wonders whether that level of SDD will end up with the downsides of both worlds.
Sources
- Birgitta Böckeler (Thoughtworks), "Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl", martinfowler.com, 15 October 2025.
- GitHub spec-kit, "Specification-Driven Development (SDD)" (no date shown; accessed on 10 October 2026).
- Kent Beck, "Augmented Coding: Beyond the Vibes", Software Design: Tidy First?, 25 June 2025.
- Cucumber, "Behaviour-Driven Development" (accessed on 10 October 2026).

Bernat López is founder and CEO of onext, an AI boutique. He helps development and product teams work with AI with method —specification before building, a person who decides where there is risk and Spec-Driven Development— and applies to his own company what he proposes: onext runs on its own agentic system.
LinkedIn →