Coffee value · validation and local serving pilot · 23 September 2026

JEV improves price RMSLE and speeds up extraction

The selected JEV-feature price model lowers historical validation RMSLE by 2.14% versus the incumbent. On eight public product pages, complete JEV extraction took 0.423 seconds at the median versus 4.879 seconds for the previous OpenAI extractor. The local pilot excludes page fetch and prediction, and the validation gain still needs a fresh test set.

−2.14%price RMSLE, best JEV hybrid vs incumbent
0.423 smedian complete JEV extraction, 16 calls
−91.3%complete extraction median vs previous app extractor

Complete extraction timing

MeasurePrevious OpenAI extractorComplete JEV extractor
Median4.879 s0.423 s
p958.061 s0.791 s
Successful calls16/1616/16

Both paths received the same freshly fetched HTML and page context. Page fetch and prediction were excluded. Each provider ran twice on eight pages in randomized order. On the six individual coffees, both paths returned the same listed price in all 12 pairs and package weights differed by less than 0.001 g. JEV declined both rotating subscriptions in both rounds; the previous extractor appraised them as individual coffees. This small local sample does not establish production end-to-end latency or general extraction accuracy. Inspect the paired timings and field outcomes.

Prediction performance

Task and metricIncumbentJEV A+B hybridBest tested follow-upInterpretation
Rating concordance ↑0.8910440.892528—+0.001484; below the proposed +0.005 target
Rating MAE ↓1.0501.047—Small improvement
Price RMSLE ↓0.2591660.2543540.253610 with Panama × GeshaBest gain: 2.14%; below the proposed 3% target
Price MAE, USD/100 g ↓3.8753.9123.914MAE rises by about 1%

The added exact-lot auction question yielded RMSLE 0.254796, slightly worse than the original JEV hybrid’s 0.254354. The Panama × Gesha term uses two existing JEV answers and requires no extra provider question. The most expensive price decile remains strongly underpredicted. Exploratory paired intervals cross zero for the rating gain [-0.0015, +0.0050] and the best price candidate’s RMSLE change [-0.0129, +0.0025]. These historical validation rows were repeatedly used during research, so the gains need a fresh evaluation. See the rating and original JEV report and price follow-up detail.

Earlier semantic-call pilot

MeasureCurrent app extractor
gpt-4o-mini
JEV A+B
jev-1.13.0
Observed reduction
p504.81 s0.42 s91.2%
p9511.50 s0.91 s92.1%
Successful calls16/1616/16JEV faster in all 16 pairs

Eight product pages across three roasters, two uncached calls per provider per frozen page context, one call at a time in randomized order on the same local host. The page fetch and context construction were done once before timing. p95 is the maximum of only 16 observations and is unstable. The JEV input-token charge estimate was $0.0093 for 16 successful calls; OpenAI usage was not recorded, so no cost comparison is claimed.

Frozen pageContext charactersApp median, sJEV median, s
Gesha Salma Bermudez Finca El Paraisohydrangea.coffee3,0429.150.44
Gesha Washed CRD Elida Estate, Centro Loma Lothydrangea.coffee2,3715.850.40
Hydrangea Dropshydrangea.coffee5,5134.220.49
The Originalblackwhiteroasters.com2,1394.500.39
The New Schoolblackwhiteroasters.com3,2104.270.33
G e o m e t r yonyxcoffeelab.com60,0474.770.79
Roaster's Choiceonyxcoffeelab.com7,1404.790.33
K e n y a K a m u n y a k a A Aonyxcoffeelab.com60,0477.880.64

Scope of the timing: the app’s existing call returns a full PageExtraction, including price, package size, display tasting notes, and page classification. The JEV call returns 38 typed semantic answers for model features. Both saw the same saved page context, but These measurements predate the complete JEV app path. They establish semantic-call latency, not an end-to-end serving speedup.

Extraction quality and serving decision

All final benchmark calls completed after the shared TypeSafe client was given an application User-Agent. Two TypeSafe preflight calls received HTTP 403 before that fix and are excluded from latency statistics. On two pages whose saved contexts contain no “natural” process claim, JEV nevertheless marked natural processing supported in both rounds. A rotating subscription page also drew inconsistent attributes from historical offerings in the app extractor. The new complete extraction comparison above addresses product identity on this small page set; conflicting process claims still need review.

The earlier semantic-call pilot returned fewer fields than the app. The complete-extraction comparison above is the relevant serving-component measurement. A fresh price-model test and production end-to-end latency comparison remain outstanding. The lightweight rating model is unchanged.

Inspect the public evidence file for the aggregate model metrics, page-context hashes, and individual call timings and semantic outputs. It excludes the third-party page text and credentials. The frozen contexts and full research artifacts remain local, so an exact replay requires those files; refetching a page can produce different content. The local harness is scripts/benchmark_jev_serving_extraction.py.