Skip to content

The points process

A point is DevFellowship’s unit of work. It is what a fellow is paid on, what a client is charged on, and what a client removes when they remove scope. One number, three jobs — which is exactly why it is worth knowing where it comes from.

This page describes the process as it is implemented today, not as it is planned. Where a piece is missing, it says so.

Two things assign a point value, and only two:

WhoWhereReviewed?
The Spec Builder generator (an LLM)dfl-ai-spec-builder-n8n-proxy Edge Function → work.ai_spec_tasks.estimated_pointsNo gate. Generated items are written immediately.
A human — usually the founderwork.tasks.points, edited directlyIt is the review.

Nothing else does. There is no scoring service, no decision table, no model behind an API. Every automatically-assigned point in the system was produced by one LLM call against one prose rubric, and that rubric is the whole engine.

On 2026-07-09 the generator turned one spec into 23 tasks totalling 122 points. The founder reviewed them and corrected the epic to 22.5 — a 5.42× deflation. No delivery had been issued, so the gap never became an invoice.

The error was not mysterious, and it was not the model. The rubric — inherited verbatim from a retired n8n workflow — was wrong in three compounding ways:

  1. The scale inflated. It declared a Fibonacci deck (1, 2, 3, 5, 8, 13, 21) with an hours table (1: <1h · 2: 1-2h · 3: 2-4h) implying roughly 0.5–1.2 hours per point, against a real anchor of 2 hours per point.
  2. The floor was 1. No 0, no 0.25, no 0.5 existed in the deck, so work worth nothing still scored at least a point.
  3. Fragmentation was mandated. “Uma feature completa DEVE gerar no mínimo 3-5 tarefas.” Forced fragmentation multiplied by a floor of 1 is a compounding inflator.

Splitting the correction separates the two error modes cleanly: 32% of the 122 points sat on work scored at ~zero — documentation, a duplicate, two trivial removals, two migrations — items that should not have been separate tasks at all. The remaining 14 genuine tasks went 83 → 21, a 3.95× deflation that is the scale-definition error on its own.

The rubric was retired and rewritten on 2026-08-05. Re-running the identical input through the live function three times now yields 31 · 33.5 · 32.5 (mean 32.3) — a 3.78× improvement, with the hard rules firing unprompted: all four documentation tasks scored 0, both migrations 0.5, both trivial removals 0.5. Those are the same corrections the founder had made by hand, reproduced without a human in the loop.

Everything the generator is told about assigning a number: the anchor and scale, the hard rules that pre-empt the scale, and the granularity rules that decide how many tasks the points get spread across.

Reproduced verbatim from system-prompt.ts in devfellowship/dfl-schema @ 2b93ad70a715. This block is generated and checked in CI (scripts/check-points-doc-drift.mjs); a change to the live rubric turns this page red.

### 5. Estimativa de Pontos
**Âncora: 1 ponto = 2 horas de trabalho de um dev sênior, medidas PRÉ-IA.**
A âncora é deliberadamente pré-IA: ganho de produtividade por ferramenta de IA
move a margem do negócio, não o score da tarefa.
**Escala DFL** — a esmagadora maioria das tarefas cai nestes valores:
| Pontos | Equivalente | Quando usar |
|---|---|---|
| **0** | — | Não billável (ver 5.1) |
| **0.25** | ~30 min | Ajuste pontual, remoção trivial, troca de constante ou de texto |
| **0.5** | ~1 h | A menor entrega real: código + commit + PR + review |
| **1** | ~2 h | Mudança pequena e autocontida |
| **2** | ~4 h | Tamanho mediano típico da DFL |
| **3** | ~6 h | Trabalho de quase um dia |
Valores acima de 3 são permitidos, em passos inteiros (4, 5, 6, ...), quando o
trabalho for **genuinamente** maior. **NUNCA quebre uma tarefa apenas para caber
na escala** — uma tarefa de 8 pontos é uma resposta válida e é PREFERÍVEL a oito
tarefas de 1 ponto.
**`0`, `0.25` e `0.5` são valores de primeira classe.** Se o trabalho leva meia
hora, a resposta é `0.25` — não `1`. Arredondar para cima é o erro mais caro que
este agente pode cometer.
Na dúvida entre dois valores adjacentes, escolha o **menor**.
#### 5.1 Regras duras (aplique ANTES da tabela; elas vencem a escala)
- **Documentação não é billável.** Tarefa cujo produto é documento, README,
changelog, wiki, diagrama explicativo ou apresentação → **0 pontos**.
- **Migration é barata.** Criar ou alterar uma migration de banco → **0.5 pontos**,
salvo se envolver backfill de dados ou reescrita de contrato.
- **Duplicatas colapsam.** Se dois itens entregam a mesma coisa, emita **uma única
tarefa** com os pontos dela — nunca duas tarefas repetindo os mesmos pontos em
cada uma.
### 6. Granularidade de Tarefas
**Uma tarefa = uma unidade entregável de trabalho.** NÃO existe número mínimo nem
máximo de tarefas: o input decide. Um input simples pode e deve gerar UMA tarefa.
- **NÃO** decomponha uma tarefa para aumentar a contagem, para caber na escala de
pontos, ou por hábito. Fragmentar infla a estimativa e é proibido.
- Se dois itens seriam feitos no mesmo commit, no mesmo PR ou pela mesma pessoa na
mesma sessão, eles são **UMA** tarefa.
- **NÃO** invente tarefas satélites que o trabalho já implica — "escrever testes",
"criar a migration", "documentar", "revisar", "configurar ambiente" só viram
tarefa própria se o input pedir explicitamente E forem entregáveis separados.
- Decomponha **apenas** quando as partes forem entregáveis de forma independente:
ordens diferentes, pessoas diferentes ou PRs diferentes.
  1. Prose in. A spec, a conversation, a meeting transcript — free text — is POSTed to the Edge Function by create_spec_run (or the deprecated generate_tasks), carrying the caller’s JWT, never a service identity.

  2. The hard rules are applied first. Documentation → 0. A migration → 0.5. Two items delivering the same thing collapse into one. These are stated to pre-empt the scale, because they are rules rather than judgement calls — each one is a correction the founder previously made by hand.

  3. The scale is applied to what survives. 0 · 0.25 · 0.5 · 1 · 2 · 3, with whole numbers above 3 when the work is genuinely larger. Ties break downward. Rounding up is named in the rubric as the most expensive mistake available.

  4. Granularity is decided, not assumed. There is no minimum or maximum task count. Two items that would land in the same PR are one task. An 8-point task is a valid answer and is preferred over eight 1-point tasks — fragmenting to fit the scale inflates the estimate and is forbidden.

  5. The shape is validated, the value is not. generator.ts rejects a task with an invented stage_id or a missing field. It does not cap, round, or sanity-check points — a number the model emits is the number that is stored.

  6. A human edits. The generated items are candidates. Points are corrected in place, and that edit is the only review gate that exists.

  7. Promotion into execution. An approved item’s points are copied into work.tasks.points, which is the single source of truth for every downstream money number: v_delivery_metrics.total_value multiplies summed task points by the delivery’s price_per_point, and payments.invoice_items.points flows into billing.

The score answers one question — how much work is this? — and it is intended to be client-blind: the same work scores the same regardless of who the invoice goes to. The client premium is charged once, on the rate card, not twice by also inflating the count.

┌──────────────┐
│ SCORE (pts) │ ← one number, client-blind
└──────┬───────┘
┌─────────────┴─────────────┐
▼ ▼
┌──────────────────┐ ┌───────────────────┐
│ CHARGE │ │ PAY │
│ pts × charge_rate│ │ pts × pay_rate │
│ client-scoped │ │ role/tier-scoped │
└──────────────────┘ └───────────────────┘
│
▼
SCOPE REMOVAL: Δ = Δ pts × charge_rate, on the version delta

A quote is meant to be a pure function of pinned inputs — (spec_run_id, version_number, points_total, rate_card_version, client_id) — with nothing recomputed at read time. The moment a quote cites “the points” without a version, the number stops being reproducible as soon as anyone edits anything.

What exists, and what it is worth:

SurfaceStateUsable as provenance?
work.ai_spec_tasks.metadata (jsonb)'{}' on 100% of rowsNo — the obvious slot, universally empty.
work.ai_spec_tasks → work.tasksNo foreign key, and no column of any kindNo — promotion copies the values and discards the link.
work.spec_generation_logs1 row, whose four task identifiers do not exist in work.tasksNo — the intended bridge, empty and already dangling.
work.ai_spec_tasks.status'approved' on 100%, created_at = updated_at on 100%No — rows are inserted already-approved; the pending/rejected states are dead.
public.activity_logsThe only real trail. 356 points edit events over 253 tasksPartially — details->>'source' is NULL on every task event. The schema anticipated provenance; nobody wired it.
work.entity_change_logReturns 0 rows — the view explodes details->'changed_fields', the payload uses diff/{set,unset}No — broken by shape mismatch.

Provenance is reconstructible for a minority of the corpus, by joining lower(trim(title)) against lower(trim(name)) — a fragile string join that matches 110 of 164 spec tasks — and then scanning activity_logs for a points edit. That reconstruction is clean where it applies (every divergent task has a points-edit event; every agreeing task has none, confirming the AI estimate shipped untouched on 87 of them), and it is unavailable for the ~1365 tasks with no spec-task counterpart at all.

Provenance becomes columns rather than a heuristic: points_source ∈ {engine, human, generator_legacy, imported}, plus engine_version, rule_id, cell_key, confidence. The acceptance bar is zero NULL points_source on a scored run — not “mostly populated”. Estimation itself is planned to move out of the prompt into a small reviewed decision table, with the hard rules as explicit rows carrying citable rule_ids, and anything outside a trusted cell escalating to a human instead of emitting a number.

Neither is shipped. Until they are, the rubric above is the engine, and the audit trail is public.activity_logs.

This page cannot silently drift from the rubric it documents. scripts/check-points-doc-drift.mjs runs on every pull request and on a schedule:

  • it reads system-prompt.ts from dfl-schema at live HEAD — a pinned commit would compare this page against the past and call that agreement;
  • it re-extracts the rubric and compares it byte-for-byte with the generated block above;
  • it refuses to report success on a short read, a missing anchor, a missing marker, or a rubric that has lost any of its load-bearing rules — a comparison that finds no differences because it compared nothing is the failure mode the guard exists to prevent;
  • a --self-test runs first, proving the comparator still reports drift on text that differs, so a guard that has quietly stopped biting cannot pass as green;
  • being unable to read the upstream file is a red build, never a skip.

When it goes red, run node scripts/check-points-doc-drift.mjs --write — then read the diff. The block is generated; the prose around it is not.