LoRA intent adaptation / AI & apps
Independent local comparative study
Adapting a small language model to Spanish intents.
Eighteen LoRA fits test what a small language model gains from 600 or 2,400 labeled utterances. Adaptation substantially improved the measured scores over unchanged Qwen3; its advantage over a frozen encoder remained uncertain.
- My role
- Study design / model adaptation / evaluation
- When
- October 2026
- Built with
- Python · MLX · LoRA · Qwen3-1.7B · MASSIVE · scikit-learn
Outcome
With 2,400 examples, mean LoRA macro-F1 reached 70.22% versus unchanged Qwen3's 30.35%: +39.87 percentage points, with a paired 95% interval of [37.72, 41.67]. Against the frozen encoder's 69.78%, the difference was +0.44 points [−1.64, 2.97], establishing neither superiority nor equivalence. LoRA seed scores ranged from 67.75% to 73.79%. At 600 examples the mean was 60.92%; excluding exact-text overlap left 2,894 test rows and a +39.70-point gain over Qwen3 at the larger budget.
Developed with AI assistance; I am responsible for the study design and interpretation. The explorer shows saved results, not live inference. One locale, one small quantized LLM and three fitted seeds limit generalization. Intervals condition on those fitted models; secondary intervals have no multiplicity correction. Macro-F1 retains all 60 frozen labels, including one absent from the test. The comparative plan followed known TF-IDF scores and an engineering smoke run; it is not full preregistration. Five Spanish examples are post-hoc illustrations from MASSIVE, reproduced under CC BY 4.0.

The challenge
A language model can follow an instruction without reliably choosing the right intent label. I wanted to measure how much local LoRA adaptation helps with limited Spanish training data, how it compares with smaller classifiers, and what the improvement costs in training time and inference latency.
Approach
Compare methods on the same data
I used fixed, nested samples of 600 and 2,400 utterances from MASSIVE 1.1 es-ES. All four method families use the utterance alone and the same 60 intent labels. The 2,033 development examples select hyperparameters; the 2,974 official test examples are reserved for final scoring.
Retain all three training seeds
The LoRA search crosses two data budgets, three learning rates and three seeds, giving 18 fits. A single learning rate per budget is selected by mean development macro-F1, and all three fitted seeds are retained. The four-bit Qwen3 base, tokenizer and decoding stay fixed; LoRA trains for exactly two passes through each sample. Every selected model, source, environment and data identity is frozen together before final test scoring.
Measure paired differences and seed variation
I scored all 11 frozen model instances with exact whole-label matching and counted invalid outputs as errors. Paired intervals use 2,000 bootstrap draws of normalized-text clusters. A prespecified sensitivity excludes 80 test rows that overlap sampled training or development text. The three LoRA scores remain visible separately; their arithmetic mean is not an ensemble.
Keep resource costs and negative findings visible
All 18 fits used 98.02 minutes of measured training-plus-save time on an Apple M5 Pro with 24 GiB unified memory, excluding development inference, model loading and engineering work. On the same 100 development utterances, encoder-2,400 had a 7.63 ms median versus LoRA medians of 87.76–115.86 ms. Its p95 was worse: 370.96 ms versus 103.39–176.63 ms. These single-run timings preserve the irregular encoder tail without assigning a cause.
Explore how it works
Try it yourself ↓Recorded experiment
Loading the saved comparison…
More from the project

