Skip to content

spec-builder

Turn spec prose into a spec run: a first-class, versioned, linkable entity holding candidate tasks with points, review status and a stable per-item short_ref. A run has its own URL (spec-builder.devfellowship.com/history/<run_id>), can be referenced from a plan body as {{dfl-entity:spec_run:<run_id>}}, and reaches the real backlog only through an explicit promotion.

Endpointhttps://engineering.mcp.devfellowship.com/mcp
Tools20
Backing datawork.ai_spec_inputs (the run), work.ai_spec_tasks (its items), work.spec_run_versions (its version snapshots), work.comments (entity_name spec_run / spec_run_item), work.entity_connections (the plan binding), work.tasks (only on promotion). Calls the dfl-ai-spec-builder-n8n-proxy Edge Function for generation — despite the name, no n8n is involved.

A run carries two independent counters, and they must never be merged into one:

Run version (current_version)Comment version (comment_version)
Modelsthe pack’s contentthe conversation about it
Bumped byupdate_spec_run_items, regenerate_spec_run, promote_spec_run_itemsonly handoff_spec_run_comments
Not bumped byany commentany content edit
Pinnableyes — attach_entity({ mode: "pinned", rev: <spec_run_versions.id> })no; a cursor, not an artifact

A quote is derived from a run version’s stored points_total, so pin the run whenever a plan cites its points — an unpinned reference means the number moves the moment anyone edits an item. If a comment bumped the run version, a quote would move because somebody asked a question.

Items keep their id and their short_ref across every re-generation. regenerate_spec_run therefore takes an operation list, never a fresh task list:

  • split(item_id, new_items) CONTINUES item_id — same row, same id, same short_ref, points untouched — and only births new_items, each with lineage_ref, lineage_kind: "split_from" and points_needs_review: true. A human point is never divided, copied or re-estimated across a split.
  • merge(from_ids, into_id) KEEPS into_id and soft-removes from_ids with lineage_kind: "merged_into".
  • Removal is always soft, so an older pinned version keeps resolving and an already-sent short_ref never dangles.
  • An update touching a human-set estimated_points is rejected unless the item is named in repoint. The whole call is refused rather than the field silently dropped.

regenerate_spec_run has two paths, chosen by whether you pass operations:

operations omittedoperations passed
What happensthe generator runs over spec (or the prose already stored on the run)your operation list is applied
Itemsuntouchedreconciled
Run versionnot cutone version cut, commit_message required
Returnsoutcome: "generated_proposal" — the fresh pack plus proposed_operations, a complete valid operation list keyed to the live idsthe reconciliation report

spec is the only way to update a run’s stored prose (work.ai_spec_inputs.text), and it works on both paths. On the apply path it rides inside the same all-or-nothing envelope as the items: a rejected operation list leaves the prose exactly as it was.

The generator path deliberately proposes instead of applying. The generator emits tasks with no ids, so writing them against existing rows could only guess which live item each one is — or replace the set and destroy the short_refs that already-sent comment batches point at. Edit proposed_operations into the update / split / merge / remove operations you actually mean and send it back through the apply path.

A run stores no plan binding and has no plan_slug column. The binding is the {{dfl-entity:spec_run:<run_id>}} token in the plan body; work.entity_connections is the index the plans-app derives from it. So after create_spec_run, call attach_entity on the Plans MCP with { slug, type: "spec_run", locator: <run_id> } — one link per run, never one per task.

ToolDescription
create_spec_runGenerate a pack of candidate tasks from spec prose and store it as a spec run bound to a plan_slug — no epic, project or business unit required. Creates version 1 with a points_total and a stable short_ref per item. Does NOT create real tasks. Returns the deep link plus the exact attach_entity call that puts the run in the plan’s rail.
get_spec_runRead a run live, or at an exact version_number / version_id (which returns that version’s stored JSONB snapshot). Returns items, points_total, both counters and the open comment batch. An unresolvable version is an error, never a fallback to live. Read-only — the side-effect-free way to inspect the open comment batch.
list_spec_runsFind runs by plan_slug (joined through work.entity_connections), by epic_id, by status, or by recency. Produces the run_id every other tool needs, and feeds attach_entity.
update_spec_run_itemsApply human edits to items — rename, re-point, re-stage, re-tag, approve, reject. Bumps the run version with a required commit_message. Any estimated_points set here becomes points_set_by: "human" and lands in human_edited_fields, which is what later makes regenerate_spec_run refuse to overwrite it. All-or-nothing.
set_spec_run_points_provenanceRecord how an item’s points were produced — points_source (engine / human / generator_legacy / imported), engine_version, rule_id. The only sanctioned write path for those columns: they are data, so they are never corrected by a dfl-schema migration. Does not change the number, does not flip points_set_by, and cuts no version — a provenance edit cannot move points_total. engine without an engine_version is rejected before anything is written. Idempotent, all-or-nothing.
score_spec_run_itemsPrice items with the points_v2 decision table and stamp points_source: "engine" + engine_version + rule_id in the same write. You classify (work_type, complexity, is_duplicate); the table decides. The table is 14 rules of reviewable data in devfellowship/dfl-flows-definitions — three founder rulings (documentation → 0, migration → 0.5, duplicate → 0) evaluated first, then eleven trusted cells whose values are corpus medians. Outside a trusted cell it emits no number and escalates — there is no catch-all rule, and on the historical corpus it declines ~61% of rows and ~71% of points. Client-blind: no client, project or business-unit input exists anywhere in the path. dry_run defaults to true. Cuts a run version only when a point value actually changed. Refuses a points_set_by: "human" item unless its id is in repoint. ⚠️ Accuracy is fidelity to historical human judgement (MedAE 1.00 / 69.4% within ±1pt on covered rows vs 1.50 / 49.6% for always-guess-the-median), never a claim of correctness — there is no effort ground truth in the database.
regenerate_spec_runOmit operations → the generator re-runs: it reads spec (or the prose stored on the run), persists a spec you pass to work.ai_spec_inputs.text, and returns the fresh pack plus a ready-to-send proposed_operations list — writing no item and cutting no version. Pass operations and they are applied as before: keep / update / add / remove / split / merge over existing ids, every live item named exactly once, a reconciliation report and one run-version bump.
promote_spec_run_itemsInsert status: "approved" items into work.tasks, write work_task_id / work_task_identifier back, bump the run version. dry_run defaults to true. Idempotent — skips already-promoted items — with a named outcome per item so a partial failure reads as one.
list_spec_run_packagesThe scope layers of a run — the “onion” — each with its item count, point total, and what it adds over the previous layer. The delta is what makes a negotiation possible; a single total only allows “yes” or “too expensive”. A run with no layers is not broken: it is one package nobody has split yet, and the tool says so.
upsert_spec_run_packageCreate or rename one layer, keyed by (run_id, layer). ⚠️ layer is the depth from the core — 1 is the MVP — and the package of layer N contains every item whose layer is <= N, so the nesting is true by construction. It is not the number the client reads; that is client_label.
assign_spec_run_layersAssign items to layers in one batch. Each item points at the innermost package containing it. Carries the same three-part compare-and-set as the save path (version, live item count, newest updated_at) and is refused if another writer changed the run first — nothing is written. ⚠️ package_id: null takes an item out of every package; it does not mean layer 1.
delete_spec_run_packageDelete a layer. Refused while any item still points at it. The foreign key is ON DELETE SET NULL, so a direct delete would silently un-package every item of that layer and the run would lose scope nobody removed.
list_spec_run_feature_groupsThe feature groups of a run — the unit the buyer has an opinion about (“I want the AI copilot”) — with item counts and point totals, plus the items outside every group. Those two numbers add up to the run total; on b17c6fd2 the outside bucket is 14 items and 32% of the points. ⚠️ A feature group does not price anything: the scope layer sells, and a package total is a plain sum.
upsert_spec_run_feature_groupCreate or rename one group, keyed by (run_id, code). Groups are per run — there is no shared catalog, because a group’s composition is run-specific. Without an explicit sort_order the group appends rather than tying with another, so the list never reorders itself between two reads. Dependencies between groups live in notes, as prose and on purpose.
assign_spec_run_feature_groupsAssign items to groups in one batch, with the same three-part compare-and-set as the save path — refused if another writer moved first, and nothing is written. ⚠️ feature_group_id: null places the item outside every group, which is the base of the project (CI/CD, deploy, database, handover). It is a classification, not a blank.
delete_spec_run_feature_groupDelete a group. Refused while any item still points at it. The foreign key is ON DELETE SET NULL and “outside every group” asserts something — so a direct delete would make the run claim a dozen items are project base without anyone deciding it.
comment_spec_runAdd a review comment to the run (entity_name: "spec_run") or one item ("spec_run_item"), into the open batch. Never bumps either version. batch_number and context_version are stamped server-side by a trigger, so the UI and the MCP cannot disagree about which batch a comment is in.
handoff_spec_run_commentsReturn the open batch only, formatted as a ready-to-paste agent prompt with a short_ref anchor per item, stamp handed_off_at, and bump the comment version. dry_run previews without bumping. A second call with no new comments returns an empty batch — comments are never re-sent.
delete_spec_runPERMANENTLY delete one run plus its items, version snapshots and comments. work.ai_spec_tasks and work.spec_run_versions cascade from the run row; work.comments does not (polymorphic, no FK), so this tool deletes them explicitly — which is most of why it exists, since a raw delete leaves them orphaned and unreachable. dry_run defaults to true, and confirm_title must match the run’s title to commit. Refuses while a plan still references the run (run_attached_to_plan → detach it first with detach_entity) or when any item was promoted (run_has_promoted_items). The reference check counts through work.entity_target_reference_count() (SECURITY DEFINER, integer only), not through a SELECT on work.entity_connections — that table is plan-scoped, so a filtered count would miss a binding on a plan you cannot read and delete the run anyway. reference_count can therefore exceed the plan_slugs the refusal names; the gap means “ask the plan owner to detach it”. Never deletes a work.tasks row. For a run that is merely over, set status: "discarded" instead and keep the history.
generate_tasks⚠️ DEPRECATED — use create_spec_run + promote_spec_run_items. The only tool in the DFL MCP fleet that writes unreviewed AI output straight into a real backlog: it inserts into work.tasks immediately with no review gate, no versioning and no durable candidates, and its output is invisible in the Spec Builder history. It also requires an epic under the devfellowship/Revera business units, so it cannot run for a plan-only client project at all.