Sentence-level Arabic readability assessment on a 19-level ordinal scale, built for the BAREC 2026 Shared Task, Task 1, strict track. Given an Arabic sentence, we predict its readability level from 1 to 19. The strict-track rule is "Models must be trained exclusively on the training set of the BAREC Corpus." The ranked metric is quadratic weighted kappa (QWK)
We scored QWK 85.4 / Acc 38.7 on the blind test set and ranked 1st on the strict-track leaderboard. The numbers and the full submission history are in docs/results.md.
Every number in these docs that can be recomputed is checked by make verify. The rest is listed in
docs/reproducibility.md. What data the system touched, against the
strict-track rule, is in docs/compliance.md.
A 21-member weighted ensemble of fine-tuned Arabic transformer encoders, combined with QWK-optimal threshold calibration adapted to the blind set's estimated label prior.
Four things did most of the work, in order of measured effect:
- Morphological preprocessing. CAMeL Tools
d3toksegmentation is worth about +6–7 QWK over light normalization onbert-large-arabertv2, more than any modeling change that came after it. It is not universal: AraModernBERT did better on raw text. - Objective diversity on one backbone. Five ordinal losses (regression, CORN, soft labels, weighted-kappa loss, squared EMD) on AraBERTv2 decorrelate errors better than five different backbones do. Greedy selection kept the objective variants and dropped most of the backbones.
- Capacity, until it saturates. base (135M) → large (370M) was the largest single lever (+0.4 blind QWK). large → a LoRA-tuned ALLaM-7B (~19× the parameters) added +0.045† wQWK, below the measured ±0.1–0.2 noise floor, so we excluded it under the noise-floor rule we had fixed in E8.
- Target-prior threshold calibration, the deciding step of the final submission. We estimate the blind label prior with regularized BBSE, reweight the calibration set to it, and refit the 18 cut points. It improved both leaderboard metrics at once (85.3 → 85.4 QWK, +2.8 Acc). It has a clear failure mode: running BBSE a second time on its own output took the blind score back down to 85.2.
Everything else we tried failed under the same protocol: 15 refuted levers, from seed replicas and robust aggregation to pseudo-label distillation and 5-fold partitioning, plus the 7B. All of it is in docs/negative-results.md.
† Not recomputed by make verify. Provenance for each such number is in
docs/reproducibility.md.
src/slra_st/ Library: metrics, calibration, ensembling, shift estimation, train, infer
experiments/ One directory per experiment: hypothesis, method, launchers, verdict
configs/ Ensemble member/weight configs (V4 → V8) and fitted threshold sets
artifacts/ Cached member scores, per-model metadata, training and inference logs
submissions/ Every blind submission we uploaded, plus the candidates we generated
docs/ Data card, method, evaluation protocol, results, findings, compliance
data/ Only a README; the BAREC splits are not redistributed here
The trained weights (41 GB) and the corpus are not tracked. See artifacts/README.md and data/README.md for what to fetch and where it goes.
python -m venv .venv && .venv/bin/pip install -e .
# recompute every reported number from the committed caches (no GPU, no corpus)
make verify
# to train or infer you need the corpus; fetch the BAREC splits into data/ (see data/README.md)
make preprocess # builds the raw and d3tok variants
# train one member (backbone x objective x data regime)
.venv/bin/python -m slra_st.train --model aubmindlab/bert-base-arabertv2 \
--tag arabertv2_emd_d3 --variant d3tok --objective emd --bs 32 --lr 2e-5 --epochs 8
# run the shipped 21-member ensemble over unlabeled sentences
.venv/bin/python -m slra_st.infer --input data/blind_sent.parquet --ensemble final_v8 \
--out artifacts/regenerated/prediction
# check a submission file against the required format
.venv/bin/python -m slra_st.submission artifacts/regenerated/predictionartifacts/regenerated/ is gitignored, so re-running never overwrites the submission record.
| ID | Experiment | Question | Verdict |
|---|---|---|---|
| E1 | Backbone × objective matrix | Which Arabic encoder, which text variant? | AraBERTv2 + d3tok is the anchor (~84 test QWK) |
| E3 | Objective diversity | Do loss geometries decorrelate better than backbones? | Confirmed; best train-only single model 85.55 |
| E4 | Capacity | Does more capacity fix the systematic error tail? | Confirmed, +0.4 blind; XLM-R not competitive |
| E5 | Seed diversity | Do seed replicas add ensemble diversity? | Refuted, the tail is bias, not variance |
| E6 | Greedy selection | Which subset and weights? | V4 → V5 → V6 (test 86.398) |
| E7 | All-data retraining | Is train+dev genuine signal? | +0.31 honest A/B; one member diverged |
| E8 | Combiner bake-off | Greedy, learned weights, or a fixed blend? | Fixed 50/50 two-regime blend scores highest (86.94) |
| E9 | Validated enlargement | Ship the bigger blend only if it clears the gate | +0.010 → shipped as V8 (blind 85.3) |
| E10 | Accuracy frontier | How much Acc is free at fixed QWK? | Acc 35.9 → 40.4 at flat blind QWK |
| E11 | Target-prior calibration | Can the label shift be exploited? | The largest measured gain: 85.4 / 38.7 |
| E12 | Distillation, 7B, 5-fold | What else adds information? | Nothing: the system is saturated at 85.4–85.5 |
The index with dates, artifacts and per-experiment numbers: experiments/README.md.
After the deadline we pointed the same protocol at our own system. One model gives 82.97 weighted QWK†, eight give 84.32†, all 21 give 85.48†. A single pseudo-distilled base model gives 83.98†, which is 98.2%† of the full system at 1/21 of the inference cost, and about the same as a LoRA-tuned 7B Arabic LLM (83.94†) at roughly 1/50 of its cost.
Everything we measured after training lands in the same place: extra same-family members (+0.03/model†), distilled students, the 7B (+0.045†), 5-fold partition models (0.00†), and 15 post-processing ideas (0 or negative). We read 85.4–85.5 as the information ceiling for this label set and this class of sentence-level encoders. Details in docs/findings.md.
docs/compliance.md is the audit: the strict-track rule quoted in full, a table of exactly which BAREC splits each part of the shipped system touched, the ablations that are not track-eligible and that we never submitted, our submission-rate record against CodaBench's 5-per-day cap, and the questions the published rules do not settle.
Short version: we used no external data of any kind and no blind labels, but the system uses more than the training split, and the prior calibration in E11 is transductive, in that it reads the blind set's predicted label distribution as a whole. Neither point is addressed by the published rules, so we state both rather than assume them. If the strictest reading holds, the number to fall back on is V6 at blind 85.0.
Code is MIT-licensed (LICENSE). The BAREC corpus is not redistributed here and remains governed by the shared task's terms.