Understanding Spec-Pushed-Growth: Kiro, spec-kit, and Tessl

I’ve been attempting to grasp one of many newest AI coding buzzword: Spec-driven growth (SDD). I checked out three of the instruments that label themselves as SDD instruments and tried to untangle what it means, as of now.

Definition

Like with many rising phrases on this fast-paced area, the definition of “spec-driven growth” (SDD) remains to be in flux. Right here’s what I can collect from how I’ve seen it used to date: Spec-driven growth means writing a “spec” earlier than writing code with AI (“documentation first”). The spec turns into the supply of reality for the human and the AI.

GitHub: “On this new world, sustaining software program means evolving specs. […] The lingua franca of growth strikes to the next stage, and code is the last-mile method.”

Tessl: “A growth method the place specs — not code — are the first artifact. Specs describe intent in structured, testable language, and brokers generate code to match them.”

After wanting over the usages of the time period, and a few of the instruments that declare to be implementing SDD, it appears to me that in actuality, there are a number of implementation ranges to it:

Spec-first: A nicely thought-out spec is written first, after which used within the AI-assisted growth workflow for the duty at hand.
Spec-anchored: The spec is saved even after the duty is full, to proceed utilizing it for evolution and upkeep of the respective function.
Spec-as-source: The spec is the principle supply file over time, and solely the spec is edited by the human, the human by no means touches the code.

All SDD approaches and definitions I’ve discovered are spec-first, however not all try to be spec-anchored or spec-as-source. And infrequently it’s left obscure or completely open what the spec upkeep technique over time is supposed to be.

Understanding Spec-Pushed-Growth: Kiro, spec-kit, and Tessl

What’s a spec?

The important thing query by way of definitions in fact is: What’s a spec? There doesn’t appear to be a normal definition, the closest I’ve seen to a constant definition is the comparability of a spec to a “Product Necessities Doc”.

The time period is kind of overloaded in the intervening time, right here is my try at defining what a spec is:

A spec is a structured, behavior-oriented artifact – or a set of associated artifacts – written in pure language that expresses software program performance and serves as steerage to AI coding brokers. Every variant of spec-driven growth defines their method to a spec’s construction, stage of element, and the way these artifacts are organized inside a venture.

There’s a helpful distinction to be made I believe between specs and the extra normal context paperwork for a codebase. That normal context are issues like guidelines information, or excessive stage descriptions of the product and the codebase. Some instruments name this context a reminiscence financial institution, in order that’s what I’ll use right here. These information are related throughout all AI coding periods within the codebase, whereas specs solely related to the duties that truly create or change that exact performance.

An overview diagram showing agent context files in two categories: Memory Bank (AGENTS.md, project.md, architecture.md as examples), and Specs (Story-324.md, product-search.md, a folder feature-x with files like data-model.md, plan.md as example files).

It seems to be fairly time-consuming to guage SDD instruments and approaches in a manner that will get near actual utilization. You would need to attempt them out with completely different sizes of issues, greenfield, brownfield, and actually take the time to evaluate and revise the intermediate artifacts with greater than only a cursory look. As a result of as GitHub’s weblog put up about spec-kit says: “Crucially, your function isn’t simply to steer. It’s to confirm. At every section, you mirror and refine.”

For 2 of the three instruments I attempted it additionally appears to be much more work to introduce them into an current codebase, due to this fact making it even more durable to guage their usefulness for brownfield codebases. Till I hear utilization stories from individuals utilizing them for a time period on a “actual” codebase, I nonetheless have a variety of open questions on how this works in actual life.

That being stated – let’s get into three of those instruments. I’ll share an outline of how they work first (or fairly how I believe they work), and can hold my observations and questions for the top. Observe that these instruments are very quick evolving, so they could have already modified since I used them in September.

Kiro

Kiro is the only (or most light-weight) one of many three I attempted. It appears to be principally spec-first, all of the examples I’ve discovered use it for a job, or a consumer story, with no point out of the way to use the necessities doc in a spec-anchored manner over time, throughout a number of duties.

Workflow: Necessities → Design → Duties

Every workflow step is represented by one markdown doc, and Kiro guides you thru these 3 workflow steps within its VS Code primarily based distribution.

Necessities: Structured as a listing of necessities, the place every requirement represents a “Consumer Story” (in “As a…” format) with acceptance standards (in “GIVEN… WHEN… THEN…” format)

A screenshot of a Kiro requirements document

Design: In my try, the design doc consisted of the sections seen within the screenshot under. I solely have the outcomes of one among my makes an attempt nonetheless, so I’m undecided if this can be a constant construction, or if it modifications relying on the duty.

A screenshot of a Kiro design document, showing a component architecture diagram, and then collapsed sections titled Data Flow, Data Models, Error Handling, Testing Strategy, Implementation Approach, Migration Strategy

Duties: A listing of duties that hint again to the requirement numbers, and that get some additional UI components to run duties one after the other, and evaluate modifications per job.

A screenshot of a Kiro tasks document, showing a task with UI elements “Task in progress”, “View changes” next to them. Each task is a bullet list of TODOs, and ends with a list of requirement numbers (1.1, 1.2, 1.3)

Kiro additionally has the idea of a reminiscence financial institution, they name it “steering”. Its contents are versatile, and their workflow doesn’t appear to depend on any particular information being there (I made my utilization makes an attempt earlier than I even found the steering part). The default topology created by Kiro whenever you ask it to generate steering paperwork is product.md, construction.md, tech.md.

A version of the earlier overview diagram, this time specific to Kiro: The memory bank has 3 files in a steering folder called product.md, tech.md, structure.md, and the specs box shows a folder called category-label-enhancement (the name of my test feature) that contains requirements.md, design.md, tasks.md

Spec-kit

Spec-kit is GitHub’s model of SDD. It’s distributed as a CLI that may create workspace setups for a variety of frequent coding assistants. As soon as that construction is about up, you work together with spec-kit by way of slash instructions in your coding assistant. As a result of all of its artifacts are put proper into your workspace, that is probably the most customizable one of many three instruments mentioned right here.

Screenshot of VS Code showing the folder structure that spec-kit set up on the left (command files in .github/prompts, a .specify folder with subfolders memory, scripts, templates); and GitHub Copilot open on the right, where the user is in the process of typing /specify as a command

Workflow: Structure → 𝄆 Specify → Plan → Duties 𝄇

Spec-kit’s reminiscence financial institution idea is a prerequisite for the spec-driven method. They name it a structure. The structure is meant to comprise the excessive stage ideas which are “immutable” and will all the time be utilized, to each change. It’s principally a really highly effective guidelines file that’s closely utilized by the workflow.

In every of the workflow steps (specify, plan, duties), spec-kit instantiates a set of information and prompts with the assistance of a bash script and a few templates. The workflow then makes heavy use of checklists within the information, to trace essential consumer clarifications, structure violations, analysis duties, and so forth. They’re like a “definition of finished” for every workflow step (although interpreted by AI, so there isn’t any 100% assure that they are going to be revered).

A partial screenshot of the very end of the spec.md file, showing a bunch of checklists for content quality, requirement completeness, execution status.

Under is an summary for instance the file topology I noticed in spec-kit. Observe how one spec is made up of many information.

At first look, GitHub appears to be aspiring to a spec-anchored method (“That’s why we’re rethinking specs — not as static paperwork, however as residing, executable artifacts that evolve with the venture. Specs grow to be the shared supply of reality. When one thing doesn’t make sense, you return to the spec; when a venture grows advanced, you refine it; when duties really feel too massive, you break them down.”) Nevertheless, spec-kit creates a department for each spec that will get created, which appears to point that they see a spec as a residing artifact for the lifetime of a change request, not the lifetime of a function. This neighborhood dialogue is speaking about this confusion. It makes me assume that spec-kit remains to be what I might name spec-first solely, not spec-anchored over time.

Tessl Framework

(Nonetheless in non-public beta)

Like spec-kit, the Tessl Framework is distributed as a CLI that may create all of the workspace and config construction for quite a lot of coding assistants. The CLI command additionally doubles as an MCP server.

Screenshot of Cursor, showing the files Tessl created in the file tree (.tessl/framework folder), and the open MCP configuration on the right, which starts the tessl command in MCP mode

Tessl is the one one among these three instruments that explicitly aspires to a spec-anchored method, and is even exploring the spec-as-source stage of SDD. A Tessl spec can function the principle artifact that’s being maintained and edited, with the code even marked with a remark on the high saying // GENERATED FROM SPEC - DO NOT EDIT. That is at the moment a 1:1 mapping between spec and code information, i.e. one spec interprets into one file within the codebase. However Tessl remains to be in beta and they’re experimenting with completely different variations of this, so I can think about that this method may be taken on a stage the place one spec maps to a code element with a number of information. It stays to be seen what the alpha product will assist. (The Tessl crew themselves see their framework as one thing that’s extra sooner or later than their present public product, the Tessl Registry.)

Right here is an instance of a spec that I had the Tessl CLI reverse engineer (tessl doc --code ...js) from a JavaScript file in an current codebase:

A screenshot of a Tessl spec file

Tags like @generate or @take a look at appear to inform Tessl what to generate. The API part reveals the concept of defining a minimum of the interfaces that get uncovered to different components of the codebase within the spec, presumably to guarantee that these extra essential components of the generated element are absolutely underneath the management of the maintainer. Working tessl construct for this spec generates the corresponding JavaScript code file.

Placing the specs for spec-as-source at a fairly low abstraction stage, per code file, most likely reduces quantity of steps and interpretations the LLM has to do, and due to this fact the prospect of errors. Even at this low abstraction stage I’ve seen the non-determinism in motion although, once I generated code a number of occasions from the identical spec. It was an fascinating train to iterate on the spec and make it increasingly particular to extend the repeatability of the code technology. That course of jogged my memory of a few of the pitfalls and challenges of writing an unambiguous and full specification.

Observations and questions

These three instruments are all labelling themselves as implementations of spec-driven growth, however they’re fairly completely different from one another. In order that’s the very first thing to bear in mind when speaking about SDD, it isn’t only one factor.

One workflow to suit all sizes?

Kiro and spec-kit present one opinionated workflow every, however I’m fairly positive that neither of them is appropriate for almost all of actual life coding issues. Particularly, it’s not fairly clear to me how they’d cater to sufficient completely different downside sizes to be usually relevant.

Once I requested Kiro to repair a small bug (it was the identical one I used up to now to attempt Codex), it shortly grew to become clear that the workflow was like utilizing a sledgehammer to crack a nut. The necessities doc turned this small bug into 4 “consumer tales” with a complete of 16 acceptance standards, together with gems like “Consumer story: As a developer, I would like the transformation operate to deal with edge instances gracefully, in order that the system stays sturdy when new class codecs are launched.”

I had an analogous problem once I used spec-kit, I wasn’t fairly positive what measurement of downside to make use of it for. Accessible tutorials are normally primarily based on creating an software from scratch, as a result of that’s best for a tutorial. One of many use instances I ended up attempting was a function that may be a 3-5 level story on one among my previous groups. The function trusted a variety of code that was already there, it was supposed to construct an summary modal that summarised a bunch of information from an current dashboard. With the quantity of steps spec-kit took, and the quantity of markdown information it created for me to evaluate, this once more felt like overkill for the scale of the issue. It was an even bigger downside than the one I used with Kiro, but additionally a way more elaborate workflow. I by no means even completed the total implementation, however I believe in the identical time it took me to run and evaluate the spec-kit outcomes I might have applied the function with “plain” AI-assisted coding, and I might have felt far more in management.

An efficient SDD software would on the very least have to offer flexibility for a number of completely different core workflows, for various sizes and varieties of modifications.

Reviewing markdown over reviewing code?

As simply talked about, and as you may see within the description of the software above, spec-kit created a LOT of markdown information for me to evaluate. They have been repetitive, each with one another, and with the code that already existed. Some contained code already. General they have been simply very verbose and tedious to evaluate. In Kiro it was a bit simpler, as you solely get 3 information, and it’s extra intuitive to grasp the psychological mannequin of “necessities > design > duties”. Nevertheless, as talked about, Kiro additionally was manner too verbose for the small bug I used to be asking it to repair.

To be trustworthy, I’d fairly evaluate code than all these markdown information. An efficient SDD software must present an excellent spec evaluate expertise.

False sense of management?

Even with all of those information and templates and prompts and workflows and checklists, I steadily noticed the agent finally not comply with all of the directions. Sure, the context home windows at the moment are bigger, which is usually talked about as one of many enablers of spec-driven growth. However simply because the home windows are bigger, doesn’t imply that AI will correctly decide up on every part that’s in there.

For instance: Spec-kit has a analysis step someplace throughout planning, and it did a variety of analysis on the present code and what’s already there, which was nice as a result of I requested it so as to add a function that constructed on high of current code. However finally the agent ignored the notes that these have been descriptions of current courses, it simply took them as a brand new specification and generated them over again, creating duplicates. However I didn’t solely see examples of ignoring directions, I additionally noticed the agent go manner overboard as a result of it was too eagerly following directions (e.g. one of many structure articles).

The previous has proven that the easiest way for us to remain accountable for what we’re constructing are small, iterative steps, so I’m very skeptical that a number of up-front spec design is a good suggestion, particularly when it’s overly verbose. An efficient SDD software must cater to an iterative method, however small work packages virtually appear counter to the concept of SDD.

How you can successfully separate practical from technical spec?

It’s a frequent concept in SDD to be intentional in regards to the separation between practical spec and technical implementation. The underlying aspiration I assume is that finally, we might have AI fill in all of the solutioning and particulars, and change to completely different tech stacks with the identical spec.

In actuality, once I was attempting spec-kit, I steadily bought confused when to remain on the practical stage, and when it was time so as to add technical particulars. The tutorial and documentation additionally weren’t fairly per it, there appear to be completely different interpretations of what “purely practical” actually means. And once I assume again on the various, many consumer tales I’ve learn in my profession that weren’t correctly separating necessities from implementation, I don’t assume we now have an excellent observe document as a occupation to do that nicely.

Who’s the goal consumer?

Most of the demos and tutorials for spec-driven growth instruments embody issues like defining product and have objectives, they even incorporate phrases like “consumer story”. The concept right here may be to make use of AI as an enabler for cross-skilling, and have builders take part extra closely in necessities evaluation? Or have builders pair with product individuals once they work on this workflow? None of that is made specific although, it’s introduced as a given {that a} developer would do all this evaluation.

Through which case I might ask myself once more, what downside measurement and sort is SDD meant for? Most likely not for big options which are nonetheless very unclear, as absolutely that may require extra specialist product and necessities abilities, and plenty of different steps like analysis and stakeholder involvement?

A 2x2 matrix, x-axis “Clarity of problem”, y-axis “Size of problem”. Each quadrant has a box with a question mark, and there is a label in the middle that says “Where does SDD sit?”

Spec-anchored and spec-as-source: Are we studying from the previous?

Whereas many individuals draw analogies between SDD and TDD or BDD, I believe one other vital parallel to have a look at for spec-as-source particularly is MDD (model-driven growth). I labored on a number of tasks initially of my profession that closely used MDD, and I saved being reminded about that once I was attempting out the Tessl Framework. The fashions in MDD have been principally the specs, albeit not in pure language, however expressed in e.g. customized UML or a textual DSL. We constructed customized code turbines to show these specs into code.

Example of a structured, parseable specification DSL from my past experience, mostly recreated from memory. Screen “Write Message” instantiates InputScreen { … } Illustrates things like references to domain model fields, inheritance from other screens for reusability of patterns, navigation logic.

In the end, MDD by no means took off for enterprise functions, it sits at an ungainly abstraction stage and simply creates an excessive amount of overhead and constraints. However LLMs take a few of the overhead and constraints of MDD away, so there’s a new hope that we are able to now lastly give attention to writing specs and simply generate code from them. With LLMs, we’re not constrained by a predefined and parseable spec language anymore, and we don’t must construct elaborate code turbines. The value for that’s LLMs’ non-determinism in fact. And the parseable construction additionally had upsides that we’re dropping now: We might present the spec writer with a variety of software assist to jot down legitimate, full and constant specs. I’m wondering if spec-as-source, and even spec-anchoring, may find yourself with the downsides of each MDD and LLMs: Inflexibility and non-determinism.

To be clear, I’m not nostalgic about my MDD expertise up to now and saying “we’d as nicely convey that again”. However we should always look to code-from-spec makes an attempt up to now to study from them after we discover spec-driven in the present day.

Conclusions

In my private utilization of AI-assisted coding, I additionally typically spend time on rigorously crafting some type of spec first to present to the coding agent. So the overall precept of spec-first is certainly precious in lots of conditions, and the completely different approaches of the way to construction that spec are very wanted. They’re among the many high most steadily requested questions I hear in the intervening time from practitioners: “How do I construction my reminiscence financial institution?”, “How do I write an excellent specification and design doc for AI?”.

However the time period “spec-driven growth” isn’t very nicely outlined but, and it’s already semantically subtle. I’ve even just lately heard individuals use “spec” principally as a synonym for “detailed immediate”.

Concerning the instruments I’ve tried, I’ve listed a lot of my questions on their actual world usefulness right here. I’m wondering if a few of them are attempting to feed AI brokers with our current workflows too actually, finally amplifying current challenges like evaluate overload and hallucinations. Particularly with the extra elaborate approaches that create a number of information, I can’t assist however consider the German compound phrase “Verschlimmbesserung”: Are we making one thing worse within the try of creating it higher?

Understanding Spec-Pushed-Growth: Kiro, spec-kit, and Tessl

Definition

What’s a spec?

Kiro

Spec-kit

Tessl Framework

Observations and questions

One workflow to suit all sizes?

Reviewing markdown over reviewing code?

False sense of management?

How you can successfully separate practical from technical spec?

Who’s the goal consumer?

Spec-anchored and spec-as-source: Are we studying from the previous?

Conclusions

Related Articles

When You Have RSUs, ISOs, NQSOs, and an ESPP: The best way to Coordinate Fairness Compensation

How you can discover low-cost flights wherever

Ebook Evaluate: Ideas of Bitcoin

LEAVE A REPLY Cancel reply

Latest Articles

When You Have RSUs, ISOs, NQSOs, and an ESPP: The best way to Coordinate Fairness Compensation

How you can discover low-cost flights wherever

Ebook Evaluate: Ideas of Bitcoin

World well being chief is optimistic about Trump support agenda : NPR

MAN v FAT associate with Northampton Saints