feat: add Claude API extraction (internal/extract)

Stufe 1 from the core principle: Client calls the Claude Messages API
directly over net/http (no SDK dependency, stays consistent with
"Go-Standard-Library wo möglich") and forces tool-use with a strict JSON
schema instead of parsing free text. Extract() returns rules.Facts
directly rather than an intermediate DTO, since producing exactly that
is the point of this stage. Platform is supplied by the caller, never
guessed by the model.

Every failure mode returns an error instead of a zero-value Facts:
network errors, non-200 API responses, a missing tool_use block, and —
critically — a gegenleistung value outside the four allowed enum
values, which would otherwise get silently coerced into a wrong fact.
Tested entirely against an httptest fake server, no real API calls.

Building this surfaced a real gap in internal/rules: Evaluate() treated
Consideration=="unklar" the same as any other value, i.e. it just
produced an empty finding list — indistinguishable from "everything's
fine". That contradicts the core principle (uncertain extraction should
trigger a user clarification, never a judgment). Evaluate() now returns
(findings, needsClarification), with a golden case covering it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
noroot
2026-08-27 13:43:01 +02:00
parent bec34b5988
commit da91c3e2c1
6 changed files with 531 additions and 7 deletions

View File

@@ -12,8 +12,17 @@ type Finding struct {
// Evaluate prüft alle Regeln gegen f und liefert ein Finding für jede
// zutreffende Regel. Reihenfolge folgt der Reihenfolge von rules.
func Evaluate(rules []Rule, f Facts) []Finding {
var findings []Finding
//
// Ist f.Consideration "unklar", wird KEINE Bewertung abgegeben —
// needsClarification ist dann true und findings ist immer leer. Das ist
// Absicht (Kernprinzip): eine unsichere Extraktion erzeugt eine
// Rückfrage an den Nutzer, niemals eine stille "keine Findings"-
// Bewertung, die wie "alles in Ordnung" aussähe.
func Evaluate(rules []Rule, f Facts) (findings []Finding, needsClarification bool) {
if f.Consideration == ConsiderationUnclear {
return nil, true
}
for _, r := range rules {
if r.Condition.Matches(f) {
findings = append(findings, Finding{
@@ -26,5 +35,5 @@ func Evaluate(rules []Rule, f Facts) []Finding {
})
}
}
return findings
return findings, false
}

View File

@@ -16,9 +16,10 @@ import (
// dafür liefern muss. Das ist das eigentliche Asset des Projekts, nicht
// die UI — siehe CLAUDE.md.
type goldenCase struct {
Name string `json:"name"`
Facts rules.Facts `json:"facts"`
ExpectedFindings []goldenFinding `json:"expected_findings"`
Name string `json:"name"`
Facts rules.Facts `json:"facts"`
ExpectedFindings []goldenFinding `json:"expected_findings"`
ExpectNeedsClarification bool `json:"expect_needs_clarification"`
}
type goldenFinding struct {
@@ -60,7 +61,11 @@ func TestGolden(t *testing.T) {
t.Fatalf("parse %s: %v", entry.Name(), err)
}
got := rules.Evaluate(ruleSet, gc.Facts)
got, needsClarification := rules.Evaluate(ruleSet, gc.Facts)
if needsClarification != gc.ExpectNeedsClarification {
t.Fatalf("%s: needsClarification = %v, want %v", gc.Name, needsClarification, gc.ExpectNeedsClarification)
}
gotKeys := make([]string, 0, len(got))
for _, f := range got {