← All work

LLM evidence position / AI & apps

Controlled model evaluation

Testing where a language model uses evidence.

A 1,600-answer experiment moves the same supporting passage through English and Spanish contexts. At the larger input ceiling, middle-position accuracy was lower in both languages; the difference between their position effects remained uncertain.

My role
Evaluation design / local inference / analysis
When
October 2026
Built with
Python · MLX · Qwen3-1.7B · Belebele · Matplotlib

Outcome

At the larger input ceiling, accuracy with the evidence at the end versus the middle was 63% versus 53% in English and 59% versus 47% in Spanish. Paired middle-minus-last differences were −10 and −12 percentage points, with 95% passage-bootstrap intervals of [−17.6, −2.0] and [−20.6, −3.1]. The language-by-position contrast did not establish a clear difference between languages.

One quantized model and 100 paired questions support a limited comparison. Intervals are descriptive and not adjusted for multiple contrasts. Matched translated passages have unequal token lengths, so language effects cannot be separated from length and translation effects. The explorer shows recorded answers; it does not run live inference.

Measured English and Spanish accuracy at first, middle and last evidence positions under two input ceilings, with passage bootstrap intervals
Belebele · 100 paired questions from 96 passages; bars show 95% passage-bootstrap intervals. Input ceilings do not imply equal token lengths across languages. · Open full size ↗

The challenge

Receiving the correct passage does not guarantee a correct answer. I wanted a paired comparison of how a small language model uses supplied evidence at the beginning, middle and end of a context, and whether that pattern differs between English and Spanish.

Approach

Freeze a paired sample before evaluation

I selected 100 parallel Belebele questions from 96 passages with a fixed seed, balanced the correct-option slots, and preserved the same option permutation across languages. The protocol, sample, prompt tokens, source code, model hashes and environment were frozen before the first benchmark response.

Move evidence while preserving its context

Each question uses first, middle and last evidence positions under 2,048- and 4,096-token input ceilings. Whole translated passages and distractor order remain matched. No-context and evidence-only controls run once per question and language. The evidence is supplied deliberately, so this evaluates evidence use rather than retrieval.

Record every constrained answer

A pinned four-bit Qwen3-1.7B model ran locally with MLX, thinking disabled, greedy A/B/C/D selection and a fresh context cache for every case. An append-only journal preserved all 1,600 answers and supported an exact resumption after the first 32 operational checks. No model weights were trained or fine-tuned.

Compare questions and preserve uncertainty

I measured accuracy and paired position differences, then resampled whole passages in a 2,000-draw bootstrap so questions sharing a passage stayed together. Prespecified language-by-position contrasts, all controls, actual token lengths and post-hoc error examples are included in the report.

Sources & project context

Explore how it works

Try it yourself ↓

Recorded experiment

Loading the recorded results…

More from the project

Within-question middle-minus-first and middle-minus-last accuracy differences in English and Spanish, with passage bootstrap intervals
Paired position effects · differences in percentage points; intervals are descriptive and are not adjusted for multiple contrasts · Open full size ↗
English and Spanish accuracy without a passage and with only the correct evidence passage
Controls · no-context and evidence-only responses run once per question and language · Open full size ↗
Next projectCNN robustness & calibration ↗