The project in detail

Understanding Lingua Nostra

From corpus building to publishing an open model: the roadmap, the scientific method, the needs, the funding, and the questions you are asking.

A phase of scientific preparation

Lingua Nostra is today a research project at the concept stage. No prototype exists yet, and that is the logical consequence of where we currently stand.

« Notre priorité du moment est la constitution du corpus et la recherche des premiers partenaires scientifiques et financiers. »

Before building a reliable artificial intelligence for a regional language, it is essential to assemble a quality corpus, obtain the necessary permissions, and bring together the scientific and technical resources needed. That preparatory work is precisely what Lingua Nostra is engaged in right now. The absence of a demonstration is therefore not a shortcoming: it marks the starting point of a rigorous undertaking.

The corpus, concretely

Already identified and in the public domain:

Still to be surveyed and negotiated:

Assembling this corpus is the first phase of the project, not an established fact.

Six phases, from collection to transmission

The project advances by milestones, not by dates. Each phase prepares the next. Here is the planned progression.

1
Corpus building

Gathering and digitising Provençal texts: Félibrige almanacs and press, reference dictionaries, Provençal literature. Optical character recognition on older editions. This is the longest and most decisive step.

2
Cleaning and normalisation

Correcting recognition errors, removing duplicates to prevent the model from rote-learning, balancing the Provençal varieties as much as possible, and formatting the whole into a standard suitable for training.

3
Model training

Starting from a large multilingual model and adapting it to Provençal in three steps: continued pre-training, where the model absorbs the written language; instruction fine-tuning, where it learns to follow instructions rather than merely continue text, using thousands of question-and-answer pairs drafted and corrected by native speakers; then alignment to their preferences, which anchors register, accuracy and spelling norms.

4
Linguistic validation

A network of native speakers validates, corrects and chooses between two answers, through a simple interface designed for a non-technical audience. Without them, no quality model is possible: at this stage they are the only available source of truth.

5
Publication

Releasing the model and the code under an open licence, self-hostable, together with corpus documentation. The tool then belongs to the language community, not to any institution.

6
Extension to other languages

Transposing the method proven on Provençal to other languages of France. The second language targeted would be Breton, in particular its Vannetais variety. In the longer term, other languages of the Midi could each receive their own model — Gascon, Languedocian, Limousin — alongside other regional languages. Lou Pichoun serves as the laboratory; those that follow benefit from everything it will have taught us.

The logic behind the method

Why a corpus is indispensable. In natural language processing (NLP), a large language model learns by observing vast quantities of text. For it to speak Provençal, it must be exposed to Provençal in sufficient quantity and quality. The value of the final artificial intelligence model depends directly on that of the corpus: it is the raw material of the entire project.

Why adapting an existing large model rather than starting from scratch. Training a language model from nothing requires volumes of data and computing power far beyond the reach of a project like this one. A large multilingual model has already learned the general structure of language; it then only needs to be reoriented toward Provençal. Several low-resource languages, from Turkish to Irish, have been successfully adapted this way with modest means, even without a closely related high-resource language to draw on.

The three stages of training. First, continued pre-training, where the model absorbs the written language and learns to extend its sentences naturally. Then instruction fine-tuning, a stage common to every language assistant: the model learns to follow instructions using thousands of question-and-answer pairs covering translation, grammar, culture and conversation. What is specific to Provençal is that these pairs are drafted and corrected one by one by native speakers, making this the most demanding human contribution to the project. Finally, preference alignment, where its style is refined by showing it, between two answers, which one a speaker prefers.

How validation works. The model is evaluated separately in each dialect. During fine-tuning, it systematically produces two answers; a Rhodanian or Maritime speaker selects the best one in their own dialect. Each dialect is thus validated by its own speakers: there is no single Provençal variety to arbitrate, but two lines refined in parallel.

How performance is measured. Several instruments complement one another: perplexity, which measures the model's ease with unseen text; fill-in-the-blank tests, where a word is masked to check whether the model recovers it; a question-and-answer set on Provençal culture and literature; and above all an assessment by several independent experts judging grammaticality, linguistic accuracy and register naturalness.

The long-term ambition: a free Provençal voice. Natural language processing does not stop at text. In time, Lou Pichoun aims to add an open-source Provençal speech synthesis system (Text-to-Speech, TTS), trained on recordings of consenting native speakers. The challenge is not merely technical: generating a voice is today a solved problem. What matters is the governance of that voice — informed consent, open weights, a voice belonging to the community, quality on a par with proprietary solutions. It is this ethical and sovereign framework that will make the Provençal voice a lasting resource for Provence, rather than an export product. Automatic speech recognition (transcription) would complete the loop: understanding, responding, and reading aloud in Provençal.

What remains to be met

A research project has unknowns. Stating them clearly is part of what makes an approach serious.

Writing in two dialects. The model will need to produce equally well in Rhodanian and Maritime Provençal. This is a real challenge, doubling the corpus, validation and evaluation work.

The quality of the generated Provençal. Achieving an irreproachable standard of language remains the main unknown; only speaker validation will resolve it.

A challenge already under control: mobilising speakers. The associative and Félibrige network makes it straightforward to bring together the speakers needed to correct and fine-tune the model.

What the project requires

Funding Lingua Nostra means making specific resources possible. This is not yet a detailed budget, but an account of what is needed to move from concept to the first phase.

Documentary collection

Surveying Provençal texts, identifying dormant holdings at universities and associations, locating rights holders.

Digitisation

Scanning works not yet digitised, applying optical character recognition to older editions, proofreading after machine processing.

Annotation and validation

Mobilising and supporting the speakers who will correct the model's outputs and choose between two versions.

Computing

Renting the computing power needed for training, on specialised servers, for the duration of the intensive phases.

Linguistic expertise

Drawing on expertise in Provençal and sociolinguistics to ensure rigour of language and dialectal choices.

Software development

Building the training pipeline, the community validation interface and the self-hostable final tool.

The funding model

Lingua Nostra's funding rests on several successive sources. A first crowdfunding round launches the project; institutional, private and academic funding follows.

The primary purpose of the funds is to compensate the founder's working time for the project's start-up phase, followed immediately by the first corpus collection expenses, technical costs and grant application work.

The initial phase: crowdfunding. A pre-sale crowdfunding campaign will open the way with a launch target of €5,000. Rewards will take the form, in priority, of beta access to the model and acknowledgement in the project documentation, supplemented by a few Provençal items.

€10
Named acknowledgement on the site and in the corpus documentation.
€30
Credit in the contributor roll of the official documentation.
€75
Early access to the Lou Pichoun model in beta.
€150
Provençal mug, beta access and credit in the contributor roll.
€250
Provençal t-shirt and beta access.
€350
Provençal t-shirt and cap, plus credit in the contributor roll.
€500
Provençal set (t-shirt, cap, mug) and beta access, with credit in the contributor roll.
€800
Full set, priority beta access and prominent named acknowledgement as a founding patron in the official documentation.

Next: institutional, private and academic funding. A complete application file has already been assembled to pursue these sources of support. Once the start-up phase is funded, Lingua Nostra will take the form of a French research and development company dedicated to natural language processing, which opens access to several schemes: innovative start-up status, research and innovation tax credits, a company-based doctoral convention, national and European programmes for endangered and regional languages of France, corporate and foundation patronage, and computing credits offered by model providers. Crowdfunding is therefore only the first step of a multi-tier funding strategy.

Full transparency: rewards exceeding one quarter of the contribution legally reclassify the collection as a pre-sale rather than a donation. Funds collected by a company are taxable income, including the portion paid as remuneration. The project assumes this fully: it is real income for real work.

Be part of the first circle

The project is in its early stages. Leave your address to follow its progress and be notified when Lou Pichoun launches.

By leaving your address, you agree to receive news about the project. Unsubscribe at any time.

What the project builds on

Adapting a large model to a low-resource language is not a hypothesis: several teams have done it, sometimes with very little data and modest hardware. Here are the works that mark the way.

Models for minority languages

Irish Gaelic · 2025

A bilingual Irish-English model and the project's primary methodological reference: continued pre-training followed by fine-tuning, validated by native speakers. Approximately 44 hours of computation on two graphics processors.

Irish Gaelic · 2024

The Irish pioneer: continued pre-training with vocabulary extension. The reference for adaptation methods.

Basque · 2024

Adaptation to Basque, a language structurally more distant from French than Provençal. Evaluated on real language examinations.

Turkish · 2024

Adaptation on near-consumer hardware, starting from roughly 274 million words. The most directly reproducible case.

Sámi · 2024

Usable results from a tiny corpus, comparable to the available Provençal corpus. Proof that a small volume is enough to get started.

Traditional Chinese, Portuguese · 2024-2025

Two demonstrations of economic feasibility: a complete pipeline on a single consumer graphics card, and a full training run for approximately $500.

Resources and projects for the languages of France

Inria · 2023-2027

A national programme of corpora and tools for the languages of France (Alsatian, Breton and others). Provençal is not yet represented: Lingua Nostra would fill that gap.

Armana prouvençau · Lou Tresor dóu Felibrige
Provençal corpus · public domain

The Félibrige almanac and Mistral's great dictionary, digitised and freely available. The foundation of the Provençal corpus.

What people ask us

Why start with Provençal?

Because Provençal is seriously endangered: this project is first and foremost an act of preservation. Because Provençal, legible to French speakers from the region, is a lever for language reclamation. Because it is the language of the project's founder, who has access to the speaker network needed for validation. And because no model speaks Provençal today: this would be a first.

Why not simply use ChatGPT?

Large commercial models do not master Provençal: the language is almost entirely absent from their training data. When prompted, they produce an approximate Provençal mixed with French. They can neither speak it properly nor serve as a judge of what is correct.

Why build a specialised model?

A dedicated model, trained on a Provençal corpus validated by native speakers, reaches a quality that no general-purpose model offers for Provençal. It can be self-hosted, published under an open licence, and remain under the community's control rather than depending on a proprietary service.

How much data is needed?

The realistic target is between 50 and 100 million words of Provençal. That may seem modest, but the Northern Sámi model produced usable results with four times less. The available Provençal corpus falls within this viable range.

Why self-hosting?

For technological independence and data sovereignty. A self-hosted model depends on no company, cannot be withdrawn or modified by a third party, and guarantees the longevity of the tool and control over the data of the speakers who made it possible.

Why an open model?

So that the tool belongs to the language community rather than to any institution. An open licence allows anyone to use, study, improve and redistribute the model. That is the condition for genuine digital sovereignty of the language.

Why is there no prototype yet?

Because the quality of a model depends first of all on the quality of its corpus. Before any training can begin, that corpus must be assembled, permissions obtained, and scientific and technical resources secured. That preparatory phase is where the project stands today: a premature prototype would tell us nothing reliable.

Back to the overview