LLM evidence position / AI & apps
Controlled model evaluation
Testing where a language model uses evidence.
A 1,600-answer experiment moves the same supporting passage through English and Spanish contexts. At the larger input ceiling, middle-position accuracy was lower in both languages; the difference between their position effects remained uncertain.
- My role
- Evaluation design / local inference / analysis
- When
- October 2026
- Built with
- Python · MLX · Qwen3-1.7B · Belebele · Matplotlib
Outcome
At the larger input ceiling, accuracy with the evidence at the end versus the middle was 63% versus 53% in English and 59% versus 47% in Spanish. Paired middle-minus-last differences were −10 and −12 percentage points, with 95% passage-bootstrap intervals of [−17.6, −2.0] and [−20.6, −3.1]. The language-by-position contrast did not establish a clear difference between languages.
One quantized model and 100 paired questions support a limited comparison. Intervals are descriptive and not adjusted for multiple contrasts. Matched translated passages have unequal token lengths, so language effects cannot be separated from length and translation effects. The explorer shows recorded answers; it does not run live inference.

The challenge
Receiving the correct passage does not guarantee a correct answer. I wanted a paired comparison of how a small language model uses supplied evidence at the beginning, middle and end of a context, and whether that pattern differs between English and Spanish.
Approach
Freeze a paired sample before evaluation
I selected 100 parallel Belebele questions from 96 passages with a fixed seed, balanced the correct-option slots, and preserved the same option permutation across languages. The protocol, sample, prompt tokens, source code, model hashes and environment were frozen before the first benchmark response.
Move evidence while preserving its context
Each question uses first, middle and last evidence positions under 2,048- and 4,096-token input ceilings. Whole translated passages and distractor order remain matched. No-context and evidence-only controls run once per question and language. The evidence is supplied deliberately, so this evaluates evidence use rather than retrieval.
Record every constrained answer
A pinned four-bit Qwen3-1.7B model ran locally with MLX, thinking disabled, greedy A/B/C/D selection and a fresh context cache for every case. An append-only journal preserved all 1,600 answers and supported an exact resumption after the first 32 operational checks. No model weights were trained or fine-tuned.
Compare questions and preserve uncertainty
I measured accuracy and paired position differences, then resampled whole passages in a 2,000-draw bootstrap so questions sharing a passage stayed together. Prespecified language-by-position contrasts, all controls, actual token lengths and post-hoc error examples are included in the report.
Explore how it works
Try it yourself ↓Recorded experiment
Loading the recorded results…
More from the project

