Technology
Built on the same research as the rest of Maltese NLP, not apart from it
NeuroMaltese isn't a black box wrapped around a generic multilingual model. Here's what's actually happening underneath, and the research it's grounded in.
What's different technically
The Maltese-specific problems most models don't bother solving
Semitic-Romance hybrid grammar
Maltese fuses Semitic root-and-pattern morphology with Romance and English vocabulary, a combination no other language shares, which is why models trained on Romance or Semitic languages alone transfer poorly.
Low-resource training approach
Maltese makes up roughly 0.03% of Common Crawl. Rather than hoping scale compensates, NeuroMaltese is trained and fine-tuned on curated Maltese-specific corpora and evaluated against Maltese-specific benchmarks.
Code-switching architecture
Maltese-English code-switching is modelled as a first-class case, not a fallback, since it is how Maltese speakers actually communicate day to day.
Evaluation methodology
Performance is measured with metrics built or validated for low-resource language pairs, including COMET-based evaluation for English-Maltese, rather than metrics developed for high-resource pairs and assumed to transfer.
How it's built
How the models are trained and evaluated
- 01
Data sourcing
Maltese-specific corpora are collected and curated from parliamentary records, news broadcasts, and everyday speech and text, rather than filtered out of a general web crawl.
- 02
Tokenisation and fine-tuning
Tokenisation choices are tested specifically for Maltese, since the wrong tokeniser can quietly cost accuracy before training even starts, and models are fine-tuned rather than prompted zero-shot.
- 03
Terminology and domain adaptation
Legal, government, and financial registers are treated as distinct fine-tuning targets, not one generic Maltese output.
- 04
Benchmark evaluation
Models are evaluated against Maltese-specific test sets and metrics, including COMET-based scoring for translation, rather than metrics built for high-resource languages.
Published, peer-reviewed
This isn't internal marketing copy
Co-founder Kurt Abela's research on Maltese machine translation and evaluation is published in the same venues used by the wider NLP research community.
COMET-based evaluation method for English-Maltese translation, peer-reviewed at LREC-COLING
Hybrid term-injection research for low-resource Maltese terminology translation (LoResMT)
Where this sits academically
Built alongside, not instead of, Maltese NLP research
The University of Malta's Maltese Language Resource Server (MLRS) publishes foundational Maltese corpora, lexicons, and NLP tools, and MELABench benchmarks dozens of large language models against smaller fine-tuned models specifically on Maltese tasks.
NeuroMaltese comes out of the same academic community. Kurt Abela's peer-reviewed research sits alongside this work, not apart from it, and exists to take it into production rather than compete with it for the same territory. For the wider picture of what that looks like in deployment, see Neural AI's Maltese-language AI services page.
Want the technical detail on your specific use case?
Happy to walk through benchmarks, architecture, or what training data would be needed for your domain.