Exact-lot auction JEV price comparison

Status: complete. Auction coverage: 6661/6661 price rows. The original A+B answers are reused.

Decision: All JEV hybrid point estimates improve on the incumbent. The best candidate, JEV plus Panama × Gesha, lowers RMSLE by 0.005556 (2.14%). The original JEV hybrid improves it by 1.86%. The auction feature does not add to that gain: it changes RMSLE by +0.000443 versus the original JEV hybrid, while the combined model changes it by -0.000365. The best result remains short of the predeclared 3% target, and its exploratory paired interval includes zero, so it is a provisional improvement rather than promotion evidence.

RunRMSLE ↓Gain vs incumbentRelative gainMAE ($)p90 AE ($)Top-decile mean prediction ($)Top-decile bias ($)Val − train RMSLE
incumbent0.259166+0.000000+0.00%3.8755.09934.26-21.95+0.086120
scrubbed control0.259126+0.000040+0.02%3.8745.09034.29-21.92+0.086039
jev original0.254354+0.004812+1.86%3.9124.96634.45-21.76+0.081830
jev auction0.254796+0.004370+1.69%3.9244.96934.33-21.88+0.082510
jev pg0.253610+0.005556+2.14%3.9145.02235.62-20.59+0.082642
jev auction pg0.253989+0.005177+2.00%3.9244.92735.50-20.71+0.083262

Price is real USD per 100 g. The top decile is selected by true validation price; negative bias means underprediction. The Panama × Gesha term is the product of two existing A+B JEV support probabilities. The auction term is a seven-outcome probability distribution for this exact reviewed lot.

Auction choices across price train and validation: auction_component_only: 2, auction_mentioned_unclear: 3, exact_lot_auctioned: 45, explicit_non_auction: 5, not_stated: 6487, other_lot_or_history: 119. Exact-lot positives: 40 training, 5 validation. Estimated TypeSafe spend for this one-question extraction: $0.2405 in successful-response charges; retries and failures may add charges.

Selected coefficients on the model's log-price scale:
jev_auction: exact-lot auction +0.2333
jev_pg: Panama × Gesha +0.6189
jev_auction_pg: exact-lot auction +0.2239, Panama × Gesha +0.6140

Exploratory paired intervals

jev_auction minus original JEV RMSLE: [-0.000168, +0.001130]
jev_pg minus original JEV RMSLE: [-0.004259, +0.002973]
jev_auction_pg minus original JEV RMSLE: [-0.003719, +0.003147]. These roaster-cluster intervals use historical validation rows that were repeatedly used in earlier research, so they are not a fresh confirmatory test.

Serving extraction speed

In a local eight-page, 16-pair pilot, the current app extraction call took 4.81 s at p50 and 11.50 s at p95; the A+B JEV semantic call took 0.42 s and 0.91 s. JEV does not yet return the app's price, package, and display fields, so this is a call-latency pilot, not a complete serving speedup. See the combined results and method.

The final serving report should compare identical saved page contexts with cache bypassed and equivalent successful outputs. Report extraction p50 and p95, end-to-end appraisal p50 and p95, paired per-page differences, completion rate, fallbacks, and request cost. The proposed gates—30% lower extraction p50, 20% lower extraction p95, and 20% lower end-to-end p50 with no p95 regression—are reasonable directional targets; measured before/after values should remain visible even when a gate is missed.

The targeted 61-row development pilot separated exact auction lots from generic auction-system history, blends containing an auction component, a different record auction, explicit non-auction lots, and no-auction controls. Row 3497, named “Special Auction,” remains ambiguous because the source does not directly say it was auctioned.

Manual review of all 45 exact-lot positives found four without an explicit auction claim in the supplied text (row IDs 2798, 2864, 4084, and 4836). Row 3497 is ambiguous. The validation split has only 5 exact-lot positives, whose mean actual price is $19.19 per 100 g; the 40 training positives average $81.79. This small, shifted subgroup limits what the auction comparison can establish.

The baseline is the saved research price artifact; other runs refit the same fixed ElasticNet architecture and scrub explicit target quotations. Review prose may differ from roaster product pages. No serving model or inference path was changed.