# Scoring methodology: EUDR legality document assessment **Working spec, version 0.3 · August 2026** Applies to the TRACT EUDR legality analyser. Reference implementation: `lib/analyse.mjs`, `lib/countryRisk.mjs`, `lib/reference.mjs`. ## Which methodology document to use There are two editions of this methodology. They describe the same method. | Edition | File | Use it for | |---|---|---| | **Published document** | [EUDR Legality Document Assessment Methodology, v0.2](docs/TRACT_EUDR_Legality_Assessment_Methodology_v0.2.pdf) | Sharing outside the team: customers, prospects, auditors, colleagues. Designed, paginated, TRACT branded. | | **Web edition** | [docs/methodology-web-v0.3.html](docs/methodology-web-v0.3.html), published as a shareable link | Sending a URL rather than an attachment. Same content as this file. | | **Working spec** | This file, v0.3 | Engineering reference. Tracks the current build and moves ahead of the published PDF between releases. | **The published PDF is at v0.2 and the working spec is at v0.3.** Three things in this file are not yet in the PDF: verification steps and corroboration requests now carry an impact rating; calibration moved to a server-side example endpoint, so it can no longer be requested for an uploaded document; and example results are pre-computed. The published document should be regenerated at v0.3 before the next external share. The core method is unchanged in both: same five components, same weights, same bands, same aggregation. Changes in 0.3: house writing style applied to all generated text (plain English, no em dashes, short sentences). Calibration mode moved to a server-side sample endpoint. Example results are pre-computed. Verification steps and corroboration requests now carry an impact rating. No change to components, weights, bands or aggregation. --- ## 1. What the score is, and what it is not The tool produces a single **assessed confidence score (0 to 100)** for a document submitted as evidence of legality under Regulation (EU) 2023/1115. **It answers this question:** how much weight should a careful reviewer place on this document as Article 9(1)(h) evidence, given what is visible on its face and what is known about the origin? **It does not answer:** is this document genuine? Nothing is checked against an issuing register. For most producer countries, no public register exists to check. A high score means the instrument is of a strong type, well formed, from a plausible authority, in a context where it can be corroborated. It does not mean the document has been verified. The distinction is load-bearing. Output must never state or imply that a document has been verified, confirmed genuine, or validated. Every assessment carries an explicit `verification_status` sentence. It records that no register was contacted. **The score is an input to human review. It is not a compliance determination.** --- ## 2. Legal basis | Provision | Role in the method | |---|---| | **Art. 9(1)(h)** | The obligation being evidenced: "adequately conclusive and verifiable information" that commodities were produced in accordance with the relevant legislation, including any arrangement conferring the right to use the area. The `evidential_sufficiency` component measures against this. | | **Art. 2(40)(a) to (h)** | Defines "relevant legislation" as eight categories: land use rights; environmental protection; forest-related rules; third parties' rights; labour rights; human rights under international law; FPIC; tax, anti-corruption, trade and customs. Pass 1 maps every document against all eight. | | **Art. 10(2)(g)** | "The source, reliability, validity, and links to other available documentation" of the Art. 9(1) information. This is the basis for the four document-level components. | | **Art. 10(2)(h)** | Country-level concerns: corruption, prevalence of document and data falsification, weak law enforcement, human rights violations, armed conflict, sanctions. This is the basis for the `country_context` component. | | **Art. 10(2)(n)** | Certification schemes are complementary information only. An ISPO or MSPO certificate is scored as corroboration. It is never a substitute for Art. 9 information. | | **Art. 10(3)** | A valid FLEGT licence creates a presumption of compliance with Art. 3(b) for in-scope wood. The scoring recognises this. Note it covers the legality limb only, not the deforestation-free limb. | | **Art. 29** | Country benchmarking tiers. An input to `country_context`. | --- ## 3. Pipeline Two model passes run in sequence. Pass 2 depends on Pass 1's output, so they cannot run in parallel. ``` Document (PDF / JPG / PNG) │ ├─► PASS 1 Classification │ Inputs: the document, plus the origin reference library │ Outputs: document type, issuing authority, origin, fields read from the face, │ coverage of all 8 Art. 2(40) categories, Art. 9(1) data points supplied, │ observations, legibility │ └─► PASS 2 Validity assessment Inputs: the document, Pass 1 output, and the country risk profile Outputs: 5 component scores, red flags, verification steps, corroboration needed, verification status │ ▼ DETERMINISTIC AGGREGATION (code, not model) │ ▼ Overall score and band ``` Both passes use Claude Opus 5 with adaptive thinking at high effort and JSON-schema-constrained structured output. The document itself is re-sent in Pass 2, so scoring is grounded in the artefact rather than in Pass 1's description of it. All generated text follows a fixed style rule set in the prompts: plain English, short sentences, roughly an 8th-grade reading level, exact legal citations, and no em dashes. --- ## 4. The five components The model scores each component from 0 to 100 and writes a finding that cites specifics from the document. ### 4.1 Document integrity (weight 0.20) Is the document complete and properly executed on its face? Expected fields present. Signatures and seals where the instrument requires them. All pages supplied. Legible. No signs of alteration. Not a partial extract presented as a whole. *Lowers the score:* missing signature blocks, absent seals, a single page of a multi-page instrument, referenced annexes not attached, poor legibility, evidence of tampering. ### 4.2 Internal consistency (weight 0.20) Do the document's own contents agree with each other? Dates in sequence. Areas and percentages that add up. The same party named consistently. Identifiers in the correct format. The issuing authority matching the instrument type. Arithmetic that holds. *This is the component most likely to surface a fabrication.* In testing, it independently caught that a specimen's declared Reserva Legal plus APP exceeded its declared remaining native vegetation. That implies an undisclosed Forest Code deficit. The prompt did not specify that relationship. ### 4.3 Issuing authority plausibility (weight 0.20) Does the named issuing body exist? Could it lawfully issue this instrument for this place? Checked against the origin reference library: the right authority for the document type, the right administrative level, and territorial competence over the stated location. *An instrument issued by a body that could not lawfully issue it is the strongest single indicator of a problem.* Example: a land certificate over Indonesian forest estate that has never been released from it. BPN cannot lawfully title such land, so the certificate's existence is itself the red flag. ### 4.4 Evidential sufficiency for Art. 9(1)(h) (weight 0.25) How far does this document actually go towards "adequately conclusive and verifiable information"? **The instrument is judged for what it is, not for what it is labelled.** The governing hierarchy: | Strength | Instrument class | Examples | |---|---|---| | Strong | Registered title with a searchable chain | Matrícula (BR), Land Title Certificate (GH), HGU/SHM (ID), LURC (VN) | | Moderate | Statutory registration or licence with legal effect | Certificat Foncier Rural (CI), MPOB licence (MY), STDB (ID) | | Weak | Customary or unregistered grant | Attestation villageoise (CI), allocation note without concurrence (GH), girik/SKT (ID) | | Weak | Self-declared registry entry | CAR (BR) | | Complementary only | Third-party certification | ISPO, MSPO, voluntary schemes, per Art. 10(2)(n) | This component carries the highest weight because it speaks most directly to the statutory test. ### 4.5 Country context, Art. 10(2)(h) (weight 0.15) How much does the operating environment reduce the weight this document can carry? **This component scores the environment, not the document.** A flawless document from a low-verifiability, high-falsification origin still scores lower here. Inputs: - **Art. 29 benchmarking tier** (low, standard or high) - **Corruption Perceptions Index 2025** score - **Curated country and sector concerns**: documented falsification prevalence, enforcement gaps, tenure-system characteristics, sector-specific risks - **Registry verifiability**: whether documents of this kind can be independently checked at all Verifiability weighs heavily within this component. In practice it matters more than the benchmarking tier. Brazil (standard tier, CPI 35) scores above Vietnam (low tier, CPI 40) on country context. The reason: Brazilian documents carry codes designed for public online validation, while Vietnamese land records have no public register at all. --- ## 5. Aggregation **The overall score is computed in code, not by the model.** ``` overall = round( Σ (componentScore × componentWeight) ) clamped to [0, 100] ``` Weights sum to exactly 1.00. A reviewer can reproduce the number by hand from the five published component scores. That is the point. ### Bands | Score | Band | Meaning | |---|---|---| | 80 to 100 | **Substantial** | Can carry significant weight in the due diligence file, subject to the verification steps. | | 60 to 79 | **Moderate** | Useful evidence, but not sufficient on its own for the categories it covers. | | 40 to 59 | **Limited** | Weak evidence. Needs corroboration before it can support a negligible-risk conclusion. | | 0 to 39 | **Insufficient** | Should not be relied on as legality evidence in its current form. | The band is derived from the computed score. It is never supplied independently. ### Worked example Vietnam Land Use Right Certificate, calibration run: | Component | Score | Weight | Contribution | |---|---:|---:|---:| | Document integrity | 66 | 0.20 | 13.2 | | Internal consistency | 70 | 0.20 | 14.0 | | Issuing authority plausibility | 86 | 0.20 | 17.2 | | Evidential sufficiency | 70 | 0.25 | 17.5 | | Country context | 56 | 0.15 | 8.4 | | | | | **70.3 → 70 (moderate)** | --- ## 6. The determinism boundary Deciding what belongs to the model and what belongs to code is the central design choice. | Owned by **code** | Owned by the **model** | |---|---| | Weights and their sum | Component scores | | The weighted-sum arithmetic | Written findings justifying each score | | Band thresholds and derivation | Red flags | | 0 to 100 clamping | Verification steps and their feasibility tags | | Component deduplication and gap-filling | Corroboration requirements | | Canonical component ordering | Overall rationale and decisive factor | **Why the arithmetic moved into code.** In development, the model was asked to compute the weighted overall itself. On a test document it returned 26 where its own component scores implied 34. Both fell in the same band, so the conclusion was unaffected. But a score a reviewer cannot reproduce is not defensible in a compliance file. The model now supplies judgement only. **Robustness handling.** The model has been observed to emit the same component twice. Post-processing keeps the first occurrence of each, restores canonical order, and fills any omitted component with a zero score and an explicit "not assessed" finding. A malformed response degrades visibly rather than silently skewing the total. **Schema constraint.** Structured outputs reject `minimum` and `maximum` on integer fields. The 0 to 100 range is stated in the field description and enforced by clamping after parsing. --- ## 7. Calibration guidance given to the model Anchors are specified in the prompt to prevent clustering: - A well-executed, internally consistent **registered title from a moderate-risk origin sits at 75 to 85, not 95.** The residual reflects that nothing has been verified. - A **self-declared registry printout offered as tenure evidence belongs at 25 to 40**, however cleanly it is printed. - **Above 85 is reserved** for instruments that are both intrinsically strong and independently checkable. - The model is told not to cluster scores in the middle, and to distinguish clearly between strong and weak instruments. **Specimen rule.** A document watermarked as a specimen, sample, template or draft, or declaring no legal effect, is decisive. `document_integrity` and `evidential_sufficiency` score very low, a high-severity red flag is raised, and the rationale says so plainly. --- ## 8. Outputs beyond the score The score alone is not actionable. Each assessment also returns: **Red flags.** Specific signs that the document may not be what it claims to be, or may not support what it is being used to support. Severity low, medium or high, each with why it matters. The model is instructed not to manufacture flags. **Verification steps.** Ranked, concrete next actions, each tagged by feasibility: | Tag | Meaning | |---|---| | `public_online` | Checkable now, free, without the supplier | | `published` | In a published list or register | | `in_country_request` | Requires a formal application in-country | | `supplier_dependent` | Cannot proceed without the supplier's cooperation | | `not_verifiable` | No independent source exists | Each step and each corroboration request also carries an **impact rating** (high, medium or low). High means it addresses the decisive factor, or could confirm or refute the document itself. The model is told most assessments have only one or two high-impact steps. The UI sorts and groups by these ratings, so a reviewer sees the needle-movers first. Steps must be specific. "Verify the title" is rejected. "Run a SICAR consultation on CAR number MG-3143906 and compare the returned polygon area against the 148.7320 ha declared" is the standard. Steps that do not depend on the supplier come first. **Corroboration needed.** The documents that should sit alongside this one to close the Art. 9(1)(h) gap for the categories it touches. **Verification status.** A plain sentence recording that no issuing register was contacted, and what that means for reading the score. --- ## 9. Country risk data Held in `lib/countryRisk.mjs` for six origins: cocoa in Côte d'Ivoire and Ghana, coffee in Brazil and Vietnam, oil palm in Indonesia and Malaysia. | Field | Source | |---|---| | `eudrTier` | Commission Implementing Regulation (EU) 2025/1093 of 22 May 2025 | | `cpi` | Transparency International CPI 2025 | | `concerns` | Curated per origin, tied to specific Art. 2(40) categories | | `registryVerifiability` | Curated: an overall rating plus an itemised list of what can be checked, how, and at what access level | **Benchmarking caveat.** EUR-Lex blocks automated retrieval, so tiers were compiled from secondary reporting of the Annex. **Sources disagree on Ghana and Indonesia.** Every entry carries a `tierConfidence` field recording the disagreement, and the UI surfaces a caveat where sources conflict. Confirm at the official source before relying on a tier for a compliance decision. **Status of the list.** The European Parliament objected on 9 July 2025 (373 to 289). It criticised the methodology and the classification of major deforestation fronts as standard rather than high risk. The objection did not block the instrument, and it remains in force. The first scheduled review falls in 2026, so tiers may change. **Effect of the tier.** A low-risk tier permits Art. 13 simplified due diligence and lowers the competent authority check rate: 1% of operators, against 3% standard and 9% high. It does **not** remove the Art. 9(1) information obligation, and it does **not** lower the Art. 9(1)(h) evidential standard. The tier is therefore one input to one component at 0.15 weight, not a shortcut. When a document comes from outside the six covered origins, no curated profile exists. The model is instructed to cap `country_context` accordingly and to say so. --- ## 10. Validation to date Three synthetic specimens, each built with a known defect. Calibration mode (section 11) was used so the scoring exercised instrument strength rather than bottoming out on the specimen watermark. These are the current pre-computed example results, which are what the tool displays. | Specimen | Doc | Consist | Auth | Evid | Country | Overall | Designed defect. Detected? | |---|---:|---:|---:|---:|---:|---:|---| | Vietnam LURC | 70 | 72 | 85 | 68 | 57 | **71** moderate | Front page only; reverse-side encumbrances unknown. Yes ✅ | | Brazil CAR | 74 | 72 | 85 | 32 | 66 | **64** moderate | Self-declared registry offered as tenure evidence. Yes ✅ | | Ghana allocation note | 52 | 80 | 58 | 33 | 45 | **53** limited | No Lands Commission concurrence; incomplete chain. Yes ✅ | The Brazil row shows the components separating cleanly. It is a genuine, well-formed federal instrument (integrity 74, authority 85) that is still weak evidence for the purpose it was submitted for (evidential 32). A single-dimension score would collapse that distinction. **Run-to-run variance.** The same three specimens have been assessed three times across development. Vietnam scored 70, 68 and 71. Brazil scored 60, 68 and 64. Ghana scored 47, 48 and 53. Bands held in every run except Ghana's third, which moved from limited to the bottom of moderate territory while staying limited. Treat a score as accurate to roughly plus or minus 5, and treat the band as the meaningful signal. The published v0.2 document records the first run. **Not yet validated:** calibration against real documents. Synthetic specimens show that the components discriminate in the expected direction. They cannot show that the absolute scores are pitched correctly for production use. --- ## 11. Calibration mode and pre-computed examples A synthetic specimen declares on its face that it has no legal effect. That correctly craters two components, and it makes the fixtures useless for testing whether the scoring separates a strong instrument from a weak one. Calibration mode scores the instrument on its merits as if genuine. Everything else is unchanged. Calibration is reachable only through the server's sample endpoint, which reads the bundled specimen files from disk. The public analyse endpoint has no calibrate input. A client cannot ask for a real document to be scored as if genuine. Example results shown in the app are **pre-computed** with `bake-samples.mjs` and served from `samples/results/`. Re-bake after any prompt or scoring change, and note that baked results carry the methodology version that produced them. --- ## 12. Known limitations 1. **No authentication.** No document is checked against an issuing register. The score reflects intrinsic quality and context only. 2. **Absolute calibration unproven.** Component discrimination is demonstrated. The pitch of the absolute numbers is not yet validated against real documents. 3. **Benchmarking tiers are secondary-sourced** and disputed for Ghana and Indonesia (section 9). 4. **Six origins only.** Documents from elsewhere are assessed without a curated country profile and without an authority reference library. That materially weakens two components. 5. **Weights are a judgement, not an empirical fit.** They were set from the structure of Art. 10(2), not derived from labelled outcome data. They are a defensible starting point, to be revised once real-world review outcomes exist. 6. **Single-document scope.** Each document is scored in isolation. The tool does not yet assess whether a set of documents together satisfies Art. 9(1)(h) for a plot or consignment. That is the actual compliance question. 7. **Non-determinism.** The model may return slightly different component scores on repeated runs of the same document. Aggregation is deterministic. The inputs to it are not. --- ## 13. Change control Any change to weights, bands or component definitions changes the meaning of previously issued scores. Bump the version at the head of this document, record the change, and do not compare scores across versions without noting it. Re-bake the example results after any such change. Extension points: - **New origin**: add an entry to `ORIGINS` in `lib/reference.mjs` and to `COUNTRIES` in `lib/countryRisk.mjs`. Prompts and UI read from both. No other change is required. - **New component or weight**: `COMPONENTS` in `lib/analyse.mjs`. Weights must sum to 1.00. - **Band thresholds**: `BANDS` in `lib/analyse.mjs`. --- ## Appendix: data sources | Source | Used for | As of | |---|---|---| | Regulation (EU) 2023/1115, consolidated text 02023R1115 (EN, 26.12.2025) | All legal definitions and obligations | Dec 2025 | | Commission Implementing Regulation (EU) 2025/1093 | Art. 29 country benchmarking tiers | May 2025 | | Transparency International CPI 2025 | Corruption indicator | 2025 | | Curated origin reference library | Document types, issuing authorities, verifiability | Aug 2026 |