The project in detail
From corpus building to publishing an open model: the roadmap, the scientific method, the needs, the funding, and the questions you are asking.
Where the project stands
Lingua Nostra is today a research project at the concept stage. No prototype exists yet, and that is the logical consequence of where we currently stand.
« Notre priorité du moment est la constitution du corpus et la recherche des premiers partenaires scientifiques et financiers. »
Before building a reliable artificial intelligence for a regional language, it is essential to assemble a quality corpus, obtain the necessary permissions, and bring together the scientific and technical resources needed. That preparatory work is precisely what Lingua Nostra is engaged in right now. The absence of a demonstration is therefore not a shortcoming: it marks the starting point of a rigorous undertaking.
Already identified and in the public domain:
Still to be surveyed and negotiated:
Assembling this corpus is the first phase of the project, not an established fact.
Roadmap
The project advances by milestones, not by dates. Each phase prepares the next. Here is the planned progression.
Gathering and digitising Provençal texts: Félibrige almanacs and press, reference dictionaries, Provençal literature. Optical character recognition on older editions. This is the longest and most decisive step.
Correcting recognition errors, removing duplicates to prevent the model from rote-learning, balancing the Provençal varieties as much as possible, and formatting the whole into a standard suitable for training.
Starting from a large multilingual model and adapting it to Provençal in three steps: continued pre-training, where the model absorbs the written language; instruction fine-tuning, where it learns to follow instructions rather than merely continue text, using thousands of question-and-answer pairs drafted and corrected by native speakers; then alignment to their preferences, which anchors register, accuracy and spelling norms.
A network of native speakers validates, corrects and chooses between two answers, through a simple interface designed for a non-technical audience. Without them, no quality model is possible: at this stage they are the only available source of truth.
Releasing the model and the code under an open licence, self-hostable, together with corpus documentation. The tool then belongs to the language community, not to any institution.
Transposing the method proven on Provençal to other languages of France. The second language targeted would be Breton, in particular its Vannetais variety. In the longer term, other languages of the Midi could each receive their own model — Gascon, Languedocian, Limousin — alongside other regional languages. Lou Pichoun serves as the laboratory; those that follow benefit from everything it will have taught us.
Scientific methodology
Why a corpus is indispensable. In natural language processing (NLP), a large language model learns by observing vast quantities of text. For it to speak Provençal, it must be exposed to Provençal in sufficient quantity and quality. The value of the final artificial intelligence model depends directly on that of the corpus: it is the raw material of the entire project.
Why adapting an existing large model rather than starting from scratch. Training a language model from nothing requires volumes of data and computing power far beyond the reach of a project like this one. A large multilingual model has already learned the general structure of language; it then only needs to be reoriented toward Provençal. Several low-resource languages, from Turkish to Irish, have been successfully adapted this way with modest means, even without a closely related high-resource language to draw on.
The three stages of training. First, continued pre-training, where the model absorbs the written language and learns to extend its sentences naturally. Then instruction fine-tuning, a stage common to every language assistant: the model learns to follow instructions using thousands of question-and-answer pairs covering translation, grammar, culture and conversation. What is specific to Provençal is that these pairs are drafted and corrected one by one by native speakers, making this the most demanding human contribution to the project. Finally, preference alignment, where its style is refined by showing it, between two answers, which one a speaker prefers.
How validation works. The model is evaluated separately in each dialect. During fine-tuning, it systematically produces two answers; a Rhodanian or Maritime speaker selects the best one in their own dialect. Each dialect is thus validated by its own speakers: there is no single Provençal variety to arbitrate, but two lines refined in parallel.
How performance is measured. Several instruments complement one another: perplexity, which measures the model's ease with unseen text; fill-in-the-blank tests, where a word is masked to check whether the model recovers it; a question-and-answer set on Provençal culture and literature; and above all an assessment by several independent experts judging grammaticality, linguistic accuracy and register naturalness.
The long-term ambition: a free Provençal voice. Natural language processing does not stop at text. In time, Lou Pichoun aims to add an open-source Provençal speech synthesis system (Text-to-Speech, TTS), trained on recordings of consenting native speakers. The challenge is not merely technical: generating a voice is today a solved problem. What matters is the governance of that voice — informed consent, open weights, a voice belonging to the community, quality on a par with proprietary solutions. It is this ethical and sovereign framework that will make the Provençal voice a lasting resource for Provence, rather than an export product. Automatic speech recognition (transcription) would complete the loop: understanding, responding, and reading aloud in Provençal.
Challenges and uncertainties
A research project has unknowns. Stating them clearly is part of what makes an approach serious.
Writing in two dialects. The model will need to produce equally well in Rhodanian and Maritime Provençal. This is a real challenge, doubling the corpus, validation and evaluation work.
The quality of the generated Provençal. Achieving an irreproachable standard of language remains the main unknown; only speaker validation will resolve it.
A challenge already under control: mobilising speakers. The associative and Félibrige network makes it straightforward to bring together the speakers needed to correct and fine-tune the model.
Needs and funding
Funding Lingua Nostra means making specific resources possible. This is not yet a detailed budget, but an account of what is needed to move from concept to the first phase.
Surveying Provençal texts, identifying dormant holdings at universities and associations, locating rights holders.
Scanning works not yet digitised, applying optical character recognition to older editions, proofreading after machine processing.
Mobilising and supporting the speakers who will correct the model's outputs and choose between two versions.
Renting the computing power needed for training, on specialised servers, for the duration of the intensive phases.
Drawing on expertise in Provençal and sociolinguistics to ensure rigour of language and dialectal choices.
Building the training pipeline, the community validation interface and the self-hostable final tool.
Lingua Nostra's funding rests on several successive sources. A first crowdfunding round launches the project; institutional, private and academic funding follows.
The primary purpose of the funds is to compensate the founder's working time for the project's start-up phase, followed immediately by the first corpus collection expenses, technical costs and grant application work.
The initial phase: crowdfunding. A pre-sale crowdfunding campaign will open the way with a launch target of €5,000. Rewards will take the form, in priority, of beta access to the model and acknowledgement in the project documentation, supplemented by a few Provençal items.
Next: institutional, private and academic funding. A complete application file has already been assembled to pursue these sources of support. Once the start-up phase is funded, Lingua Nostra will take the form of a French research and development company dedicated to natural language processing, which opens access to several schemes: innovative start-up status, research and innovation tax credits, a company-based doctoral convention, national and European programmes for endangered and regional languages of France, corporate and foundation patronage, and computing credits offered by model providers. Crowdfunding is therefore only the first step of a multi-tier funding strategy.
Follow the project
The project is in its early stages. Leave your address to follow its progress and be notified when Lou Pichoun launches.
Scientific references
Adapting a large model to a low-resource language is not a hypothesis: several teams have done it, sometimes with very little data and modest hardware. Here are the works that mark the way.
Models for minority languages
A bilingual Irish-English model and the project's primary methodological reference: continued pre-training followed by fine-tuning, validated by native speakers. Approximately 44 hours of computation on two graphics processors.
The Irish pioneer: continued pre-training with vocabulary extension. The reference for adaptation methods.
Adaptation to Basque, a language structurally more distant from French than Provençal. Evaluated on real language examinations.
Adaptation on near-consumer hardware, starting from roughly 274 million words. The most directly reproducible case.
Usable results from a tiny corpus, comparable to the available Provençal corpus. Proof that a small volume is enough to get started.
Two demonstrations of economic feasibility: a complete pipeline on a single consumer graphics card, and a full training run for approximately $500.
Resources and projects for the languages of France
A national programme of corpora and tools for the languages of France (Alsatian, Breton and others). Provençal is not yet represented: Lingua Nostra would fill that gap.
The Félibrige almanac and Mistral's great dictionary, digitised and freely available. The foundation of the Provençal corpus.
Frequently asked questions
Because Provençal is seriously endangered: this project is first and foremost an act of preservation. Because Provençal, legible to French speakers from the region, is a lever for language reclamation. Because it is the language of the project's founder, who has access to the speaker network needed for validation. And because no model speaks Provençal today: this would be a first.
Large commercial models do not master Provençal: the language is almost entirely absent from their training data. When prompted, they produce an approximate Provençal mixed with French. They can neither speak it properly nor serve as a judge of what is correct.
A dedicated model, trained on a Provençal corpus validated by native speakers, reaches a quality that no general-purpose model offers for Provençal. It can be self-hosted, published under an open licence, and remain under the community's control rather than depending on a proprietary service.
The realistic target is between 50 and 100 million words of Provençal. That may seem modest, but the Northern Sámi model produced usable results with four times less. The available Provençal corpus falls within this viable range.
For technological independence and data sovereignty. A self-hosted model depends on no company, cannot be withdrawn or modified by a third party, and guarantees the longevity of the tool and control over the data of the speakers who made it possible.
So that the tool belongs to the language community rather than to any institution. An open licence allows anyone to use, study, improve and redistribute the model. That is the condition for genuine digital sovereignty of the language.
Because the quality of a model depends first of all on the quality of its corpus. Before any training can begin, that corpus must be assembled, permissions obtained, and scientific and technical resources secured. That preparatory phase is where the project stands today: a premature prototype would tell us nothing reliable.