← All work

LoRA intent adaptation / AI & apps

Independent local comparative study

Adapting a small language model to Spanish intents.

Eighteen LoRA fits test what a small language model gains from 600 or 2,400 labeled utterances. Adaptation substantially improved the measured scores over unchanged Qwen3; its advantage over a frozen encoder remained uncertain.

My role
Study design / model adaptation / evaluation
When
October 2026
Built with
Python · MLX · LoRA · Qwen3-1.7B · MASSIVE · scikit-learn

Outcome

With 2,400 examples, mean LoRA macro-F1 reached 70.22% versus unchanged Qwen3's 30.35%: +39.87 percentage points, with a paired 95% interval of [37.72, 41.67]. Against the frozen encoder's 69.78%, the difference was +0.44 points [−1.64, 2.97], establishing neither superiority nor equivalence. LoRA seed scores ranged from 67.75% to 73.79%. At 600 examples the mean was 60.92%; excluding exact-text overlap left 2,894 test rows and a +39.70-point gain over Qwen3 at the larger budget.

Developed with AI assistance; I am responsible for the study design and interpretation. The explorer shows saved results, not live inference. One locale, one small quantized LLM and three fitted seeds limit generalization. Intervals condition on those fitted models; secondary intervals have no multiplicity correction. Macro-F1 retains all 60 frozen labels, including one absent from the test. The comparative plan followed known TF-IDF scores and an engineering smoke run; it is not full preregistration. Five Spanish examples are post-hoc illustrations from MASSIVE, reproduced under CC BY 4.0.

Measured test macro-F1 for TF-IDF, a frozen multilingual encoder and LoRA at 600 and 2,400 training examples; LoRA seed points appear beside their mean, with unchanged Qwen3 as a reference
MASSIVE es-ES · original final-test figure. LoRA markers show three individual seeds and their arithmetic mean; they are not an ensemble or confidence intervals. · Open full size ↗

The challenge

A language model can follow an instruction without reliably choosing the right intent label. I wanted to measure how much local LoRA adaptation helps with limited Spanish training data, how it compares with smaller classifiers, and what the improvement costs in training time and inference latency.

Approach

Compare methods on the same data

I used fixed, nested samples of 600 and 2,400 utterances from MASSIVE 1.1 es-ES. All four method families use the utterance alone and the same 60 intent labels. The 2,033 development examples select hyperparameters; the 2,974 official test examples are reserved for final scoring.

Retain all three training seeds

The LoRA search crosses two data budgets, three learning rates and three seeds, giving 18 fits. A single learning rate per budget is selected by mean development macro-F1, and all three fitted seeds are retained. The four-bit Qwen3 base, tokenizer and decoding stay fixed; LoRA trains for exactly two passes through each sample. Every selected model, source, environment and data identity is frozen together before final test scoring.

Measure paired differences and seed variation

I scored all 11 frozen model instances with exact whole-label matching and counted invalid outputs as errors. Paired intervals use 2,000 bootstrap draws of normalized-text clusters. A prespecified sensitivity excludes 80 test rows that overlap sampled training or development text. The three LoRA scores remain visible separately; their arithmetic mean is not an ensemble.

Keep resource costs and negative findings visible

All 18 fits used 98.02 minutes of measured training-plus-save time on an Apple M5 Pro with 24 GiB unified memory, excluding development inference, model loading and engineering work. On the same 100 development utterances, encoder-2,400 had a 7.63 ms median versus LoRA medians of 87.76–115.86 ms. Its p95 was worse: 370.96 ms versus 103.39–176.63 ms. These single-run timings preserve the irregular encoder tail without assigning a cause.

Sources & project context

Explore how it works

Try it yourself ↓

Recorded experiment

Loading the saved comparison…

More from the project

Six paired macro-F1 differences between mean LoRA and its comparison methods, with both frozen-encoder intervals crossing zero
Paired differences · 95% intervals from 2,000 normalized-text-cluster bootstrap draws, conditional on the fitted models. Secondary intervals are not adjusted for multiple comparisons. · Open full size ↗
Test macro-F1 against batch-one median latency on a logarithmic scale, showing similar mean LoRA and encoder scores with different median inference times
Recorded median latency · LoRA points average three scores and three medians. The encoder-2,400 p95 was 370.96 ms versus LoRA's 103.39–176.63 ms; its median advantage does not extend to the observed tail. · Open full size ↗
Next projectLLM evidence position ↗