Status: complete. Cached feature coverage: 8911/8911 eligible rows.
| Task | Run | Primary metric | Gain over research baseline | MAE | RMSE / p90 AE |
|---|---|---|---|---|---|
| rating | baseline | 0.891044 | 0.000000 | 1.049560 | 1.557227 |
| rating | scrubbed_text_control | 0.890228 | -0.000816 | 1.053642 | 1.562105 |
| rating | hybrid | 0.892528 | 0.001484 | 1.046997 | 1.549669 |
| rating | jev_only | 0.814291 | -0.076754 | 1.583998 | 2.370701 |
| rating | app TF-IDF reference | 0.890018 | -0.001026 | 1.074512 | 1.654510 |
| price | baseline | 0.259166 | 0.000000 | 3.875031 | 5.098572 |
| price | scrubbed_text_control | 0.259126 | 0.000040 | 3.873866 | 5.090041 |
| price | hybrid | 0.254354 | 0.004812 | 3.911948 | 4.965888 |
| price | jev_only | 0.332533 | -0.073367 | 4.803225 | 5.897929 |
Rating primary metric is pairwise concordance; higher is better. Price primary metric is RMSLE in real USD per 100 g; lower is better. Gain is positive when a candidate improves. The price fit target is log(price), and RMSLE uses log1p of predicted and actual prices.
The scrubbed-text control refits the incumbent structured and text feature architecture after removing explicit target quotations from its text. It isolates this input change from replacing structured features with JEV probabilities; it is not a fourth proposed candidate.
The hybrid gains +0.001484 rating concordance and lowers price RMSLE by 1.86% against the research incumbents. Both miss the proposed +0.005 rating and 3% price gates. On matched scrubbed text, the JEV structured replacement gains +0.002300 rating concordance and lowers price RMSLE by 0.004772. The JEV-only models regress on both tasks. At this stage the incumbents were retained. A later price follow-up selected a JEV hybrid for the live app; the rating model remains unchanged.
The most expensive price decile averages $56.21/100 g. Baseline predictions average $34.26; hybrid predictions average $34.45. Hybrid high-decile RMSLE is 0.449735, versus 0.453428 for baseline. The luxury-price compression barely changes.
Exploratory paired roaster-cluster bootstrap, 300 resamples: rating hybrid concordance delta interval [-0.001529, +0.004988]; price hybrid RMSLE delta interval [-0.010389, -0.000238]. Against the matched scrubbed-text control, rating delta interval [-0.000639, +0.005885], and price RMSLE delta interval [-0.010417, -0.000216]. These are historical, repeatedly used validation rows, not a fresh confirmatory test.
Historical validation IDs have been used repeatedly for model selection. Review prose differs from roaster copy. Fresh matched-page evaluation is still needed.
The new hybrid text is scrubbed of explicit score and price quotations before fitting. The incumbent text was not scrubbed this way, so this is an additional input difference in the hybrid comparison.
Catalog: 264 ordered probability columns; contract coffee-ab-1, model jev-1.13.0, encoder probabilities-1. Fixed split file hashes and the full feature order are in training_report.json.
Final 40-row pilot: 40/40 complete; median 0.536 s, p95 0.613 s. Full backfill: 8871 successful calls; median 0.516 s, p95 0.658 s. Estimated spend across all pilot versions and full successful responses: $3.7182. The provider's published input rate is $0.042 per million tokens. Failed and retried requests may add cost beyond this estimate.
Development pilot audit: supported claims were generally grounded in sampled source text, but the provider sometimes inferred explicit negation of one processing method from another. In row 5506 it supported washed and wet-hulled from text stating only wet-hulled. These are known extraction errors, not corrected training labels. See pilot_contract3_audit.json.
Run python scripts/backfill_jev.py pilot, inspect artifacts/jev/pilot_summary.json, then python scripts/backfill_jev.py full and python scripts/evaluate_jev.py. The backfill requires TYPESAFE_API_KEY. Artifacts and cached answers are under artifacts/jev/.
Verification passed: five contract and serialization tests, Python compilation, saved bundle prediction roundtrips, and exact validation row coverage. Run python -m unittest discover -s tests -v and python -m compileall -q coffee_value/extraction scripts/backfill_jev.py scripts/evaluate_jev.py.