The points process
A point is DevFellowship’s unit of work. It is what a fellow is paid on, what a client is charged on, and what a client removes when they remove scope. One number, three jobs — which is exactly why it is worth knowing where it comes from.
This page describes the process as it is implemented today, not as it is planned. Where a piece is missing, it says so.
What awards points
Section titled “What awards points”Two things assign a point value, and only two:
| Who | Where | Reviewed? |
|---|---|---|
| The Spec Builder generator (an LLM) | dfl-ai-spec-builder-n8n-proxy Edge Function → work.ai_spec_tasks.estimated_points | No gate. Generated items are written immediately. |
| A human — usually the founder | work.tasks.points, edited directly | It is the review. |
Nothing else does. There is no scoring service, no decision table, no model behind an API. Every automatically-assigned point in the system was produced by one LLM call against one prose rubric, and that rubric is the whole engine.
Why the rubric is the leverage
Section titled “Why the rubric is the leverage”On 2026-07-09 the generator turned one spec into 23 tasks totalling 122 points. The founder reviewed them and corrected the epic to 22.5 — a 5.42× deflation. No delivery had been issued, so the gap never became an invoice.
The error was not mysterious, and it was not the model. The rubric — inherited verbatim from a retired n8n workflow — was wrong in three compounding ways:
- The scale inflated. It declared a Fibonacci deck (
1, 2, 3, 5, 8, 13, 21) with an hours table (1: <1h · 2: 1-2h · 3: 2-4h) implying roughly 0.5–1.2 hours per point, against a real anchor of 2 hours per point. - The floor was 1. No
0, no0.25, no0.5existed in the deck, so work worth nothing still scored at least a point. - Fragmentation was mandated. “Uma feature completa DEVE gerar no mínimo 3-5 tarefas.” Forced fragmentation multiplied by a floor of 1 is a compounding inflator.
Splitting the correction separates the two error modes cleanly: 32% of the 122 points sat on work scored at ~zero — documentation, a duplicate, two trivial removals, two migrations — items that should not have been separate tasks at all. The remaining 14 genuine tasks went 83 → 21, a 3.95× deflation that is the scale-definition error on its own.
The rubric was retired and rewritten on 2026-08-05. Re-running the identical
input through the live function three times now yields 31 · 33.5 · 32.5
(mean 32.3) — a 3.78× improvement, with the hard rules firing unprompted:
all four documentation tasks scored 0, both migrations 0.5, both trivial
removals 0.5. Those are the same corrections the founder had made by hand,
reproduced without a human in the loop.
The rubric, verbatim
Section titled “The rubric, verbatim”Everything the generator is told about assigning a number: the anchor and scale, the hard rules that pre-empt the scale, and the granularity rules that decide how many tasks the points get spread across.
Reproduced verbatim from
system-prompt.tsindevfellowship/dfl-schema@2b93ad70a715. This block is generated and checked in CI (scripts/check-points-doc-drift.mjs); a change to the live rubric turns this page red.
### 5. Estimativa de Pontos
**Âncora: 1 ponto = 2 horas de trabalho de um dev sênior, medidas PRÉ-IA.**A âncora é deliberadamente pré-IA: ganho de produtividade por ferramenta de IAmove a margem do negócio, não o score da tarefa.
**Escala DFL** — a esmagadora maioria das tarefas cai nestes valores:
| Pontos | Equivalente | Quando usar ||---|---|---|| **0** | — | Não billável (ver 5.1) || **0.25** | ~30 min | Ajuste pontual, remoção trivial, troca de constante ou de texto || **0.5** | ~1 h | A menor entrega real: código + commit + PR + review || **1** | ~2 h | Mudança pequena e autocontida || **2** | ~4 h | Tamanho mediano típico da DFL || **3** | ~6 h | Trabalho de quase um dia |
Valores acima de 3 são permitidos, em passos inteiros (4, 5, 6, ...), quando otrabalho for **genuinamente** maior. **NUNCA quebre uma tarefa apenas para caberna escala** — uma tarefa de 8 pontos é uma resposta válida e é PREFERÍVEL a oitotarefas de 1 ponto.
**`0`, `0.25` e `0.5` são valores de primeira classe.** Se o trabalho leva meiahora, a resposta é `0.25` — não `1`. Arredondar para cima é o erro mais caro queeste agente pode cometer.
Na dúvida entre dois valores adjacentes, escolha o **menor**.
#### 5.1 Regras duras (aplique ANTES da tabela; elas vencem a escala)
- **Documentação não é billável.** Tarefa cujo produto é documento, README, changelog, wiki, diagrama explicativo ou apresentação → **0 pontos**.- **Migration é barata.** Criar ou alterar uma migration de banco → **0.5 pontos**, salvo se envolver backfill de dados ou reescrita de contrato.- **Duplicatas colapsam.** Se dois itens entregam a mesma coisa, emita **uma única tarefa** com os pontos dela — nunca duas tarefas repetindo os mesmos pontos em cada uma.
### 6. Granularidade de Tarefas
**Uma tarefa = uma unidade entregável de trabalho.** NÃO existe número mínimo nemmáximo de tarefas: o input decide. Um input simples pode e deve gerar UMA tarefa.
- **NÃO** decomponha uma tarefa para aumentar a contagem, para caber na escala de pontos, ou por hábito. Fragmentar infla a estimativa e é proibido.- Se dois itens seriam feitos no mesmo commit, no mesmo PR ou pela mesma pessoa na mesma sessão, eles são **UMA** tarefa.- **NÃO** invente tarefas satélites que o trabalho já implica — "escrever testes", "criar a migration", "documentar", "revisar", "configurar ambiente" só viram tarefa própria se o input pedir explicitamente E forem entregáveis separados.- Decomponha **apenas** quando as partes forem entregáveis de forma independente: ordens diferentes, pessoas diferentes ou PRs diferentes.How a score becomes a number
Section titled “How a score becomes a number”-
Prose in. A spec, a conversation, a meeting transcript — free text — is POSTed to the Edge Function by
create_spec_run(or the deprecatedgenerate_tasks), carrying the caller’s JWT, never a service identity. -
The hard rules are applied first. Documentation →
0. A migration →0.5. Two items delivering the same thing collapse into one. These are stated to pre-empt the scale, because they are rules rather than judgement calls — each one is a correction the founder previously made by hand. -
The scale is applied to what survives.
0 · 0.25 · 0.5 · 1 · 2 · 3, with whole numbers above 3 when the work is genuinely larger. Ties break downward. Rounding up is named in the rubric as the most expensive mistake available. -
Granularity is decided, not assumed. There is no minimum or maximum task count. Two items that would land in the same PR are one task. An 8-point task is a valid answer and is preferred over eight 1-point tasks — fragmenting to fit the scale inflates the estimate and is forbidden.
-
The shape is validated, the value is not.
generator.tsrejects a task with an inventedstage_idor a missing field. It does not cap, round, or sanity-checkpoints— a number the model emits is the number that is stored. -
A human edits. The generated items are candidates. Points are corrected in place, and that edit is the only review gate that exists.
-
Promotion into execution. An approved item’s points are copied into
work.tasks.points, which is the single source of truth for every downstream money number:v_delivery_metrics.total_valuemultiplies summed task points by the delivery’sprice_per_point, andpayments.invoice_items.pointsflows into billing.
From points to money
Section titled “From points to money”The score answers one question — how much work is this? — and it is intended to be client-blind: the same work scores the same regardless of who the invoice goes to. The client premium is charged once, on the rate card, not twice by also inflating the count.
┌──────────────┐ │ SCORE (pts) │ ← one number, client-blind └──────┬───────┘ ┌─────────────┴─────────────┐ ▼ ▼ ┌──────────────────┐ ┌───────────────────┐ │ CHARGE │ │ PAY │ │ pts × charge_rate│ │ pts × pay_rate │ │ client-scoped │ │ role/tier-scoped │ └──────────────────┘ └───────────────────┘ │ ▼ SCOPE REMOVAL: Δ = Δ pts × charge_rate, on the version deltaA quote is meant to be a pure function of pinned inputs — (spec_run_id, version_number, points_total, rate_card_version, client_id) — with nothing
recomputed at read time. The moment a quote cites “the points” without a version,
the number stops being reproducible as soon as anyone edits anything.
Where provenance is recorded
Section titled “Where provenance is recorded”What exists, and what it is worth:
| Surface | State | Usable as provenance? |
|---|---|---|
work.ai_spec_tasks.metadata (jsonb) | '{}' on 100% of rows | No — the obvious slot, universally empty. |
work.ai_spec_tasks → work.tasks | No foreign key, and no column of any kind | No — promotion copies the values and discards the link. |
work.spec_generation_logs | 1 row, whose four task identifiers do not exist in work.tasks | No — the intended bridge, empty and already dangling. |
work.ai_spec_tasks.status | 'approved' on 100%, created_at = updated_at on 100% | No — rows are inserted already-approved; the pending/rejected states are dead. |
public.activity_logs | The only real trail. 356 points edit events over 253 tasks | Partially — details->>'source' is NULL on every task event. The schema anticipated provenance; nobody wired it. |
work.entity_change_log | Returns 0 rows — the view explodes details->'changed_fields', the payload uses diff/{set,unset} | No — broken by shape mismatch. |
Provenance is reconstructible for a minority of the corpus, by joining
lower(trim(title)) against lower(trim(name)) — a fragile string join that
matches 110 of 164 spec tasks — and then scanning activity_logs for a points
edit. That reconstruction is clean where it applies (every divergent task has
a points-edit event; every agreeing task has none, confirming the AI estimate
shipped untouched on 87 of them), and it is unavailable for the ~1365 tasks with
no spec-task counterpart at all.
What is planned
Section titled “What is planned”Provenance becomes columns rather than a heuristic: points_source ∈ {engine, human, generator_legacy, imported}, plus engine_version, rule_id,
cell_key, confidence. The acceptance bar is zero NULL points_source on
a scored run — not “mostly populated”. Estimation itself is planned to move out
of the prompt into a small reviewed decision table, with the hard rules as
explicit rows carrying citable rule_ids, and anything outside a trusted cell
escalating to a human instead of emitting a number.
Neither is shipped. Until they are, the rubric above is the engine, and the
audit trail is public.activity_logs.
Keeping this page honest
Section titled “Keeping this page honest”This page cannot silently drift from the rubric it documents. scripts/check-points-doc-drift.mjs
runs on every pull request and on a schedule:
- it reads
system-prompt.tsfromdfl-schemaat live HEAD — a pinned commit would compare this page against the past and call that agreement; - it re-extracts the rubric and compares it byte-for-byte with the generated block above;
- it refuses to report success on a short read, a missing anchor, a missing marker, or a rubric that has lost any of its load-bearing rules — a comparison that finds no differences because it compared nothing is the failure mode the guard exists to prevent;
- a
--self-testruns first, proving the comparator still reports drift on text that differs, so a guard that has quietly stopped biting cannot pass as green; - being unable to read the upstream file is a red build, never a skip.
When it goes red, run node scripts/check-points-doc-drift.mjs --write — then
read the diff. The block is generated; the prose around it is not.