Authors: Anderson Morillo, Edwin Puertas, Juan Carlos Martinez-Santos
Venue: CLEF 2026 Working Notes, CEUR Workshop Proceedings, Vol. 4283. HIPE-2026, Jena, Germany. Published 4 October 2026.
Links: PDF · CEUR-WS volume · Code · Models
Overview
Under the team name VerbaNexAI II, this paper compares two systems for HIPE-2026: classifying temporally scoped relations between person and location mentions in multilingual historical documents. One system is a fine-tuned multilingual NLI encoder (submitted as run1). The other is a compact, self-distilled language model trained on gold-filtered teacher predictions (submitted as run3). The language model also writes a short rationale for each prediction.
Problem
Digital humanities needs to know who appears where, and when, across archival collections. HIPE-2026 turns that into a per-pair classification task over OCR-noisy English, French, and German text. Each candidate pair gets a three-way at label (was the person at the location at any time before publication: TRUE, PROBABLE, or FALSE) and a binary isAt label on a short horizon around publication. Systems are ranked by macro recall averaged over both relations, which penalizes always predicting the majority class FALSE.
Method
Approach A (NLI encoder). Documents are OCR-repaired and translated into English with a distilled local LFM2.5 rewriter. Each person–place pair becomes a premise–hypothesis example for mDeBERTa-v3. TRUE maps to entailment, PROBABLE to neutral, and FALSE to contradiction. The two relations are fine-tuned separately.
Approach B (self-distilled SLM). A DSPy chain-of-thought program over Nemotron-3-Nano-4B is evolved with GEPA on Gold Train A. The optimized program generates teacher completions; only those whose predicted label matches the gold annotation are kept. The same backbone is then fine-tuned with LoRA on that filtered corpus. At inference the student is paired with a prompt evolved under GPT-OSS-20B, and malformed outputs fall back to FALSE.
Results
On the official gold labels, the self-distilled language model beats the NLI encoder on every split. Global macro recall for the submitted SLM:
- Test A English: 0.581
- Test A French: 0.588
- Test A German: 0.570
- Test B (French literary surprise): 0.534
The largest gap versus the encoder is on French Test A (+0.128 macro recall); the smallest is on English (+0.020).
On the official leaderboard (17 teams, 45 runs), VerbaNexAI II run3 placed 18th on Accuracy (0.58) and Generalization (0.57), and 12th on Accuracy–Efficiency. Run1 (the NLI encoder) placed 28th on Accuracy (0.50) and Generalization (0.44), and 17th on Accuracy–Efficiency.
Why it matters
- Shows that a 4B local model, trained only on its own correct reasoning traces, can outperform a fine-tuned NLI encoder on noisy historical text
- The SLM’s rationales give archivists an auditable reason for each person–place decision
- The efficiency ranking shows these compact local setups stay competitive when parameter count and deployed size count, not only raw accuracy