Grai · Audio data
With a clean license and consent. It is the one thing still standing between us and a Romanian model with nothing borrowed.
Our Romanian is weak for a simple, fixable reason: the speech recognition model we use saw little Romanian in training — NVIDIA's model card places Romanian in its «broad-coverage» tier, the second of three. The training plan is in progress; we will publish its cost after the first runs, not before. Money is not the blocker. Data is.
And not just any data. A corpus with the wrong license contaminates the model trained on it — if the raw material is "non-commercial", the product becomes unusable commercially. That is why we do not download whatever turns up: we check the license on each dataset's own page, at the source, and write down what we found next to it.
For substantial corpora, we buy or license, under contract. Tell us what you have and we will talk numbers.
For institutions and cultural projects: free access to what comes out of your data, with no time limit.
The data source named in the model card and on this page, if you prefer that to payment.
Segmentation, alignment, integrity filters, manifests. You provide the raw material; the rest is our job.
If a speaker withdraws consent, their segments leave the corpus. A person's voice is biometric data.
Your data stays yours. We do not redistribute it and we do not mix it into published datasets.
We checked the license on each dataset's own page. The table is useful even if you do not work with us — license mistakes in this field are expensive and common.
| Corpus | Hours | License, checked at the source | Commercial |
|---|---|---|---|
VoxPopuli ro | 89.00 measured | CC0 — public domain | yes, no strings |
Granary ro | 12,419 pseudo-labeled | CC-BY-4.0 on the annotations; the audio stays CC0 from VoxPopuli | yes |
FLEURS ro_ro | ~10 | CC-BY-4.0 | yes |
| MoRoVoc | 93 pseudo-labeled | Romanian from Romania and from Moldova, with a dialect label; MIT on the dataset card, CC-BY-4.0 in the paper — both permissive | yes |
| parlament.md | to be measured | CC BY-SA 4.0 — Romanian and Russian as spoken in Moldova, with real language switching; share-alike | conditional on share-alike, pending legal review |
| SWARA (UTCluj) | 21 | CC-BY-NC 4.0 + research agreement — non-commercial | no |
Common Voice ro | 50.6 validated | CC0, but since October 2025 only on the Mozilla Data Collective, with an account; the platform terms forbid re-hosting and re-identification | yes for internal training; no for redistribution |
WorldSpeech ro_ro + ro_md | 1,746 | CC-BY-NC-4.0 — non-commercial | no |
| RSC | 100 | CC-BY-NC-ND 4.0 — no derivatives either | no, not at all |
| Romanian or Russian telephone speech | 0 | no open dataset found | — |
| Labeled Romanian ⇄ Russian code-switching | 0 | does not exist | — |
| Our own data (calls, meetings) | ~100 | real speech, exactly the target profile | not used without legal review |
RSC is the largest clean Romanian corpus, and it is off limits. "ND" means you cannot even make derivatives — a model trained on it is a derivative. It is the most tempting and the most forbidden item on the list.
WorldSpeech has exactly what we need — 1,746 hours, including ro_md — and it
is non-commercial. The exception in Romania's Law No. 8/1996 on copyright and related rights
covers the original parliamentary recording, not their compilation. To use those hours
commercially, the extraction would have to be redone from the Romanian Senate's original archive
— weeks of work, plus a legal review we cannot give ourselves.
The balance: Romanian with a clean commercial license ≈ 173 h transcribed by humans, plus ~12,400 h automatically pseudo-labeled (Granary). Checked at the source in September 2026.
The practical conclusion: public domain and CC0 are the only ones with no strings attached. For anything else, the cleanest source is a recording made with explicit consent — which is why this page exists.
In the Republic of Moldova, Law No. 195/2024 on the protection of personal data (Legea 195/2024), has been in force since 23 August 2026. The voice of an identifiable person is biometric data — a special category, requiring an explicit legal basis, a limited purpose and a right to erasure. Any data collaboration starts there, not with a handshake.
Broadcasters, archives, cultural institutions, call centers, universities, oral history projects — tell us what you have and under what conditions it could be used. We answer every message.