NeuroMaltese By Neural AI

Built on the same research as the rest of Maltese NLP, not apart from it

NeuroMaltese isn't a black box wrapped around a generic multilingual model. Here's what's actually happening underneath, and the research it's grounded in.

The Maltese-specific problems most models don't bother solving

Semitic-Romance hybrid grammar

Maltese fuses Semitic root-and-pattern morphology with Romance and English vocabulary, a combination no other language shares, which is why models trained on Romance or Semitic languages alone transfer poorly.

Low-resource training approach

Maltese makes up roughly 0.03% of Common Crawl. Rather than hoping scale compensates, NeuroMaltese is trained and fine-tuned on curated Maltese-specific corpora and evaluated against Maltese-specific benchmarks.

Code-switching architecture

Maltese-English code-switching is modelled as a first-class case, not a fallback, since it is how Maltese speakers actually communicate day to day.

Evaluation methodology

Performance is measured with metrics built or validated for low-resource language pairs, including COMET-based evaluation for English-Maltese, rather than metrics developed for high-resource pairs and assumed to transfer.

How the models are trained and evaluated

  1. 01

    Data sourcing

    Maltese-specific corpora are collected and curated from parliamentary records, news broadcasts, and everyday speech and text, rather than filtered out of a general web crawl.

  2. 02

    Tokenisation and fine-tuning

    Tokenisation choices are tested specifically for Maltese, since the wrong tokeniser can quietly cost accuracy before training even starts, and models are fine-tuned rather than prompted zero-shot.

  3. 03

    Terminology and domain adaptation

    Legal, government, and financial registers are treated as distinct fine-tuning targets, not one generic Maltese output.

  4. 04

    Benchmark evaluation

    Models are evaluated against Maltese-specific test sets and metrics, including COMET-based scoring for translation, rather than metrics built for high-resource languages.

This isn't internal marketing copy

Co-founder Kurt Abela's research on Maltese machine translation and evaluation is published in the same venues used by the wider NLP research community.

2024

COMET-based evaluation method for English-Maltese translation, peer-reviewed at LREC-COLING

2026

Hybrid term-injection research for low-resource Maltese terminology translation (LoResMT)

Built alongside, not instead of, Maltese NLP research

The University of Malta's Maltese Language Resource Server (MLRS) publishes foundational Maltese corpora, lexicons, and NLP tools, and MELABench benchmarks dozens of large language models against smaller fine-tuned models specifically on Maltese tasks.

NeuroMaltese comes out of the same academic community. Kurt Abela's peer-reviewed research sits alongside this work, not apart from it, and exists to take it into production rather than compete with it for the same territory. For the wider picture of what that looks like in deployment, see Neural AI's Maltese-language AI services page.

Want the technical detail on your specific use case?

Happy to walk through benchmarks, architecture, or what training data would be needed for your domain.