Grai · Audio data

We are looking for
recorded Romanian speech

With a clean license and consent. It is the one thing still standing between us and a Romanian model with nothing borrowed.

Why

The data license decides the model license

Our Romanian is weak for a simple, fixable reason: the speech recognition model we use saw little Romanian in training — NVIDIA's model card places Romanian in its «broad-coverage» tier, the second of three. The training plan is in progress; we will publish its cost after the first runs, not before. Money is not the blocker. Data is.

And not just any data. A corpus with the wrong license contaminates the model trained on it — if the raw material is "non-commercial", the product becomes unusable commercially. That is why we do not download whatever turns up: we check the license on each dataset's own page, at the source, and write down what we found next to it.

What we look for

What we offer in return

Payment or licensing

For substantial corpora, we buy or license, under contract. Tell us what you have and we will talk numbers.

Access to the resulting model

For institutions and cultural projects: free access to what comes out of your data, with no time limit.

Public credit

The data source named in the model card and on this page, if you prefer that to payment.

We do the processing

Segmentation, alignment, integrity filters, manifests. You provide the raw material; the rest is our job.

Deletion on request

If a speaker withdraws consent, their segments leave the corpus. A person's voice is biometric data.

We do not resell the corpus

Your data stays yours. We do not redistribute it and we do not mix it into published datasets.

Public inventory

What already exists, and what we cannot touch

We checked the license on each dataset's own page. The table is useful even if you do not work with us — license mistakes in this field are expensive and common.

CorpusHoursLicense, checked at the sourceCommercial
VoxPopuli ro89.00 measured CC0 — public domainyes, no strings
Granary ro12,419 pseudo-labeled CC-BY-4.0 on the annotations; the audio stays CC0 from VoxPopuliyes
FLEURS ro_ro~10 CC-BY-4.0yes
MoRoVoc93 pseudo-labeled Romanian from Romania and from Moldova, with a dialect label; MIT on the dataset card, CC-BY-4.0 in the paper — both permissiveyes
parlament.mdto be measured CC BY-SA 4.0 — Romanian and Russian as spoken in Moldova, with real language switching; share-alikeconditional on share-alike, pending legal review
SWARA (UTCluj)21 CC-BY-NC 4.0 + research agreement — non-commercialno
Common Voice ro50.6 validated CC0, but since October 2025 only on the Mozilla Data Collective, with an account; the platform terms forbid re-hosting and re-identification yes for internal training; no for redistribution
WorldSpeech ro_ro + ro_md1,746 CC-BY-NC-4.0 — non-commercialno
RSC100 CC-BY-NC-ND 4.0 — no derivatives eitherno, not at all
Romanian or Russian telephone speech0 no open dataset found
Labeled Romanian ⇄ Russian code-switching0 does not exist
Our own data (calls, meetings)~100 real speech, exactly the target profilenot used without legal review

Two traps that look like opportunities

RSC is the largest clean Romanian corpus, and it is off limits. "ND" means you cannot even make derivatives — a model trained on it is a derivative. It is the most tempting and the most forbidden item on the list.

WorldSpeech has exactly what we need — 1,746 hours, including ro_md — and it is non-commercial. The exception in Romania's Law No. 8/1996 on copyright and related rights covers the original parliamentary recording, not their compilation. To use those hours commercially, the extraction would have to be redone from the Romanian Senate's original archive — weeks of work, plus a legal review we cannot give ourselves.

The balance: Romanian with a clean commercial license ≈ 173 h transcribed by humans, plus ~12,400 h automatically pseudo-labeled (Granary). Checked at the source in September 2026.

The practical conclusion: public domain and CC0 are the only ones with no strings attached. For anything else, the cleanest source is a recording made with explicit consent — which is why this page exists.

Our own rules

What we do not do with data

Applicable legal framework

In the Republic of Moldova, Law No. 195/2024 on the protection of personal data (Legea 195/2024), has been in force since 23 August 2026. The voice of an identifiable person is biometric data — a special category, requiring an explicit legal basis, a limited purpose and a right to erasure. Any data collaboration starts there, not with a handshake.

Collaboration

Do you have Romanian recordings?

Broadcasters, archives, cultural institutions, call centers, universities, oral history projects — tell us what you have and under what conditions it could be used. We answer every message.