ELRA Newsletter
Issue #11 | July 2025
Language Resources
LRs @ELRA
Language Resources in the ELRA Catalogue
Speech corpora
Comprehensive Arabic Phonetic Database
The Comprehensive Arabic Phonetic Database is a robust and detailed linguistic resource offering both phonemic and phonetic transcriptions, precisely reflecting how Modern Standard Arabic words are realized in actual speech. It is a highly comprehensive and accurate Arabic phonetic/phonemic database, covering over 329,000 entries, including over 61,000 general vocabulary entries, 101,000 Arab personal names, 143,000 foreign personal names in Arabic and 21,000 worldwide place names both Arab and non-Arab. Each entry consists of canonical forms both vocalized and unvocalized (as in natural language) accompanied by phonetic transcriptions in IPA and X-SAMPA and the user-friendly CARS phonemic transcription system. Additionally, unique features include explicit indication of vowel neutralization, accurate word stress, gender and number codes (singular or plural), and POS (part-of-speech) codes. The database is provided in a flat TSV text file.
See also the DiaLEX and ArabLEX collections for Arabic from the same provider…
EthioSpeech Corpora is comprised of over 391 hours of recorded read speech in six different Ethiopian languages by ca. 200 speakers per language: Amharic (68 hours), Tigrigna (62 hours), Oromo (70 hours), Somali (56 hours), Afar (68 hours), and Sidama (68 hours). The dominating domain is media (mainly newspapers), but for some of the languages texts from different domains were used, including spiritual contents. The recording is made using mobile devices using the LIG-Aikuma speech recording tool that is installed on the devices. The gender and age balance of readers is nearly equal for Amharic, Tigrigna and Oromo, whereas mainly male gender for the other 3 languages. The age distribution is between 18 and 40.
The ALLIES Corpus was produced within the European CHIST-Era project ALLIES (Autonomous Lifelong learnIng intelLigent Systems). It contains about 900 hours of news broadcast, including orthographic transcriptions, speaker annotations and segmentation. This corpus is based on the material that was used for the ESTER, REPERE and ETAPE evaluation packages (see ELRA Catalogue: http://catalogue.elra.info for respective packages). The ALLIES corpus was built as an extension of the previous produced corpora. It contains corrected annotations from the previous evaluation materials as well as new audio data with corresponding transcriptions. Corrections include corrected names of speakers and re-segmentation.
ISLRN submissions
The International Standard Language Resource Number (ISLRN) provides Language Resources (LRs) with unique identifiers using a standardised nomenclature. This aims to ensure that LRs are correctly identified, and consequently, recognised with proper references for their usage in applications in R&D projects, products evaluation and benchmark as well as in documents and scientific papers.
Latest figures 60 new ISLRN numbers assigned between June 2024 and May 2025.
- 60 new ISLRN numbers assigned between June 2024 and May 2025.
- A total of 3594 ISLRN numbers assigned since January 2014.
- A total of 277 distinct languages.
The latest LRs for which an ISLRN number was requested and accepted are as follows:
MATERIAL Georgian-English Language Pack
2015 NIST Language Recognition Evaluation Test Set
Comprehensive Arabic Phonetic Database
DEFT Spanish Light and Rich ERE Annotation
MATERIAL Kazakh-English Language Pack
The Xi’an Multi-Language Learner Corpus
BOLT CTS CallFriend CallHome Mainland Mandarin Chinese Audio
BOLT CTS CallFriend CallHome Mainland Mandarin Chinese Transcripts and Translations
More about ISLRN.
Legal Issues
The European Commission’s AI Continent Action Plan
The European Union has introduced the AI Continent Action Plan with the aim of establishing itself as a global leader in artificial intelligence. The plan put forward that AI development needs to align with European values and fundamental rights.
The Action Plan is structured around key pillars: building robust computing infrastructure, advancing data governance, and developing human capital through targeted skills programs. It also prioritizes the rapid adoption of AI in strategic sectors such as healthcare and public administration, supported by significant public and private investments. Another core objective is to preserve a trustworthy, human-centric AI that enhances economic competitiveness.
In the same vein, the European Commission recently launched a consultation on the use of data in Artificial Intelligence, on the simplification of the rules that apply to data and on international data flows as well as a consultation regarding high-risk AI systems.
Link: The European Commission’s AI Continent Action Plan
EDPB’s projects on upskilling and reskilling on AI and data protection
The European Data Protection Board (EDPB) continues to support the responsible development of AI through both regulatory guidance and targeted training initiatives. Following its December opinion on certain AI model development, the EDPB has launched two new projects under its Support Pool of Experts framework to address the urgent need for upskilling and reskilling in AI and data protection.
The first project, “Law & Compliance in AI Security and Data Protection”, focuses on the legal aspects of AI, providing guidance for compliance officers and legal professionals. The second, “Fundamentals of Secure AI Systems with Personal Data”, is designed for cybersecurity experts, developers, and deployers of high-risk AI systems, which put forward the technical needs to ensure secure and privacy-compliant AI solutions.
Both projects have produced detailed reports that outline essential competencies necessary to create a more favourable environment for enforcing legislation, particularly the GDPR. These reports are made available in PDF format for easy access. The EDPB has also released modifiable versions of these materials under a Creative Commons Attribution-ShareAlike license. This approach encourages the community to adapt and update the content, thus ensuring that training resources remain current and relevant.
Link 1: Law & Compliance in AI Security & Data Protection
Link 2: Fundamentals of Secure AI Systems with Personal Dat
The French DPA publishes recommendation on legitimate interest in AI development
Following a public consultation, recommendations were issued by the French data protection authority (CNIL) regarding the development of AI systems based on legitimate interest.
These recommendations define the conditions under which legitimate interest may be invoked and set out criteria relevant to practices such as web scraping. CNIL provides examples of processing activities that may be justified by legitimate interest, including the reuse of conversational data, provided that safeguards such as pseudonymization and the right to object are implemented.
Regarding web scraping, rather than prohibiting it outright, the CNIL specifies a range of measures and requirements that AI developers must observe when engaging in such activities. These include the application of clear collection criteria, the exclusion of specific data categories, and the prompt deletion of non-relevant data.
In its recommendations, CNIL instructs organizations to avoid collecting sensitive or irrelevant data and to limit data collection from particular websites, while also stressing the need for transparency and the facilitation of data subject rights.
Additionally, and in alignment with the December EDPB opinion, CNIL advises on further elements to consider when conducting the legitimate interest assessment for personal data collection via web scraping. These elements comprise establishing a list of websites excluded from collection, refraining from collecting data from sites that object to scraping, and limiting collection to data that is openly accessible, among other recommendations.
Lnk: Recommendation on AI system development based on legitimate interest
The Action Plan is structured around key pillars: building robust computing infrastructure, advancing data governance, and developing human capital through targeted skills programs. It also prioritizes the rapid adoption of AI in strategic sectors such as healthcare and public administration, supported by significant public and private investments. Another core objective is to preserve a trustworthy, human-centric AI that enhances economic competitiveness..
n the same vein, the European Commission recently launched a consultation on the use of data in Artificial Intelligence, on the simplification of the rules that apply to data and on international data flows as well as a consultation regarding high-risk AI systems.
Link: The European Commission’s AI Continent Action Plan
EDPB’s projects on upskilling and reskilling on AI and data protection
The European Data Protection Board (EDPB) continues to support the responsible development of AI through both regulatory guidance and targeted training initiatives. Following its December opinion on certain AI model development, the EDPB has launched two new projects under its Support Pool of Experts framework to address the urgent need for upskilling and reskilling in AI and data protection.
The first project, “Law & Compliance in AI Security and Data Protection”, focuses on the legal aspects of AI, providing guidance for compliance officers and legal professionals. The second, “Fundamentals of Secure AI Systems with Personal Data”, is designed for cybersecurity experts, developers, and deployers of high-risk AI systems, which put forward the technical needs to ensure secure and privacy-compliant AI solutions.
Both projects have produced detailed reports that outline essential competencies necessary to create a more favourable environment for enforcing legislation, particularly the GDPR. These reports are made available in PDF format for easy access. The EDPB has also released modifiable versions of these materials under a Creative Commons Attribution-ShareAlike license. This approach encourages the community to adapt and update the content, thus ensuring that training resources remain current and relevant.
Link 1: Law & Compliance in AI Security & Data Protection
Link 2: Fundamentals of Secure AI Systems with Personal Data
ELRA/ELDA Projects
Information on the on-going projects
Common European Language Data Space (LDS)
The Common European Language Data Space (LDS) project was launched on January 19, 2023. This 3-year project aims at establishing a European platform and marketplace for the collection, creation, sharing and re-use of multilingual and multimodal language data.
ELDA is currently working on the different tasks under its responsibility:
Setting up a technical and legal helpdesk to provide support and guidance to all platform users. The service running this helpdesk is already operational since late 2023 and it can be found here.
Setting up of a business development helpdesk which serves as a central support centre, helping data providers and consumers make the most of new opportunities in language data, artificial intelligence, and multilingual digital services. Its goal is to ensure that participants in the European Language Data Space have the tools and knowledge they need to innovate, expand, and strengthen a sustainable, multilingual digital market in Europe. The service running this helpdesk can be found here.
Definition and establishment of a Multistakeholder Data and Services Governance Scheme.
For that purpose, a solid collaboration has been established between the LDS and the Data Spaces Support Centre (DSSC) and work towards a full alignment of governance issues has been established. For instance, ELDA participates in the DSSC expert groups and in their work towards their blueprint (version 1.5 at present). Likewise, LDS collaborates closely with the Simpl and TEMS projects and regular meetings are held to evolve in a synchronised manner. Furthermore, LDS has contributed to Simpl-Open’s feasibility study on data spaces as one of their use cases. This contribution aimed at facilitating Simpl’s integration and interoperability in the data space ecosystem and, in particular, in the LDS development and deployment, also considering LDS specifications.
The governance scheme document was initially submitted in July 2024, with an updated version released in February 2025. This latest version reflects the architectural advancements of the LDS infrastructure, aligned with the release of the version 1.0.0-beta prototype in February 2025.
In addition, the work on the governance scheme is aligned with the recent establishment of the Alliance for Language Technologies EDIC (ALT-EDIC) and the key role it is expected to play in the deployment and long-term sustainability of the LDS.
Organization of events
ELDA is responsible for the organization of a large number of events in the framework of the LDS. The LDS Country Workshop series has also moved forward with the organization of the following national events, along with the official Launch Conference.
- Belgium and the Netherlands, in Breda (Leonardo Hotel) on January 23, 2025.
A joint workshop hosted by the European Language Data Space, Instituut voor de Nederlandse Taal, and Ghent University. It convened experts from Belgian and Dutch industry, public administration, and research to highlight the role of language data in advancing language technologies and AI-driven tools.
- Denmark, held in Copenhagen, Denmark (DLA Piper Oslo Plads) on March 7, 2025.
Presented by the European Language Data Space with the Alexandra Institute and the Agency for Digital Government. This workshop addressed the importance of language data in developing language technologies and AI tools, featuring panels on data production, market development, and public administration perspectives.
- Cyprus, held in Aglantzia (University of Cyprus) on 8 May 2025.
Organized with the University of Cyprus (Dept. of French & European Studies), this workshop gathered industry, public administration, and research experts to discuss language data’s critical role in supporting language technologies in Cyprus, including applications in GenAI and translation.
- Malta, held on Valletta Campus (University of Malta) on 16 May 2025.
With the University of Malta, the LDS hosted this event focused on language data’s relevance for building language technologies in Malta, it featured keynote speeches, a live LDS demonstration, and a panel on language data market creation.
- Poland, held in Warsaw (EC Representation Jasna Centre) on 29 May 2025.
Co-organized with the Institute of Computer Science of the Polish Academy of Sciences, the workshop evaluated language data’s importance for Polish AI model development. It included strategy-focused panels and presentations from government, academia, and industry.
- Estonia, held in Tallin on 11 June 2025.
Co-organised with the Institute of the Estonian Language, explored how Estonia can support its language in the age of artificial intelligence, the workshop hosted experts’ discussion on aligning EU digital strategy with national goals to keep Estonian vibrant and future-ready. Sessions highlighted the value of language data for cultural identity and economic growth. Panels focused on practical steps to develop Estonia’s language technology and data market.
- Slovenia, held in Ljubljana (Faculty of Computer and Information Science, University of Ljubljana) on 17 June 2025.
Co-organised with the Jožef Stefan Institute, University of Ljubljana, and the Slovenian Language Technologies Society, the workshop brought together experts from industry, public administration, and research. Discussions focused on the importance of language data for advancing language technologies and AI tools in Slovenia.
Technical Workshops
Workshop on Speech Recognition Solutions, held online on 1 July 2025
The European Language Data Space and the European Commission brought together leading experts from various Member States to exchange best practices, discuss and explore challenges and opportunities in Automatic Speech Recognition (ASR) solutions. The workshop aims at fostering collaboration and knowledge exchange for the benefit of all participating countries.
Ensuring the legal compliance of the Language Data Space
ELDA is currently conducting a comprehensive legal analysis of the Common European Language Data Space and evaluating it against potential governance options. This analysis encompasses a wide range of legal frameworks, including cybersecurity legislation, competition law, data protection regulations, and various EU-level digital legislation such as the Data Governance Act and Data Act. The assessment aims to ensure that the LDS governance scheme adheres to all relevant EU regulations, particularly those related to the collection and sharing of language data and services.
ALT-EDIC Inauguration, 19-21 March 2025
Cité internationale de la langue française, Villers-Cotterêts, France
The Alliance for Language Technologies (ALT-EDIC) was officially launched on March 20, 2025, by Ms Rachida Dati, the Minister of Culture and, Ms Clara Chappaz, the Minister Delegate for Artificial Intelligence and Digital Affairs of France.
To mark the occasion, several events were held, including the inaugural General Assembly of ALT-EDIC on March 19, 2025, which brought together representatives from 25 Member States, alongside stakeholders from research institutions, the developer community, and industry leaders, the LDS Launc Conference and the LLM4EU kick-off meeting.
- facilitate uptake by SMEs, NGOs, public administration, and academia of European machine translation services for websites;
- support the creation of open-source European language speech recognition solutions;
- carry out market studies on language technologies and widely disseminate their results to foster the take-up of language technologies in Europe.

LDS Launch Conference, 19 March 2025
Cité internationale de la langue française, Villers-Cotterêts, France
The LDS Launch Conference was held as part of these inaugural events on March 19, 2025 (in the afternoon). The hybrid event, held in English, provided interpretation in French and German for the online participants.
Opened by both Edouard Geoffrois, Director of the newly launched ALT-EDIC, and Philippe Gelin, Head of Sector for Multilingualism at DG CONNECT, European Commission, their introductions set the stage for insightful discussions on the collaborative foundations of the Language Data Space (LDS) initiative.
The program continued with a series of key presentations detailing the architecture and core components of the LDS ecosystem. Georg Rehm, Coordinator of the LDS, offered an overview of the marketplace developed to facilitate the sharing and monetization of language data. Stelios Piperidis (ILSP “Athena”) followed with a deep dive into the technical framework supporting the data space. Finally, Khalid Choukri (ELDA) addressed the governance model and data protection mechanisms critical to the LDS infrastructure.
See the programme: https://language-data-space.ec.europa.eu/events/lds-launch-conference-2025-03-19_en
Watch the conference (in English) on YouTube
Language Technology Solutions – CNECT/LUX/2022/OP/0030
This call for tenders from the European Commission was published within the Digital Europe programme (DIGITAL). It aims to achieve three specific goals:
- facilitate uptake by SMEs, NGOs, public administration, and academia of European machine translation services for websites;
- support the creation of open-source European language speech recognition solutions;
- carry out market studies on language technologies and widely disseminate their results to foster the take-up of language technologies in Europe.
ELRA, through its operational body ELDA, is involved in two of the funded projects which are described below.
LOT 2 –Automated speech recognition prototype solutions
The LTS Lot 2 initiative aims to develop an open-source Automatic Speech Recognition (ASR) system, and a multilingual speech dataset called LELD (Low-resource European Languages Datasets), focused on three under-resourced European languages: Czech, Estonian, and Greek. A market study on the current state of ASR technologies was also carried out.
ELDA is in charge of building the LELD corpus, while BUT and TILDE work on the ASR systems. The dataset includes 4,500 hours of speech per language, with one-third of the data transcribed (1,500 hours per language). The audio was collected from several public institutions in each country, including the European Commission, national parliaments, and local councils. The audio covers a variety of speaking styles and recording conditions, which helps train systems that are better adapted to real-world use.
Transcriptions follow rigorous guidelines and are validated through automatic processes and human cross-validation procedures to maintain high-quality outputs.
The transcribed data is used to train ASR systems for each language. Results from the first training iterations show improved recognition accuracy. All data and tools developed in the project will be made available under open licenses to support future research and development.
Consortium and Tasks
The consortium operating in this project is coordinated by Brno University of Technology (BUT) with the participation of TILDE and ELDA. Three main tasks are being performed with the participation of all members of the consortium, which are:
- Task 1: A comprehensive market study of the Automatic Speaker Recognition (ASR) solutions.
- Task 2: Creation of an open-source speech recognition prototype solution for three under-represented European languages (Czech, Estonian, and Greek).
- Task 3: Collection and partial transcription (one third) of speech data for the three above-mentioned European under-resourced languages.
ARCHER (Advancing Robust and Creative Human language technologies through CHallenge Events and Research)
ARCHER (Advancing Robust and Creative Human Language Technologies through Challenge Events and Research) is an innovative research initiative dedicated to pushing the boundaries of Automatic Speech Recognition (ASR). By leveraging state-of-the-art artificial intelligence and machine learning, ARCHER aims to enhance the accuracy, adaptability, and robustness of speech-to-text technologies across diverse environments and languages.
ARCHER’s mission is to bridge the gap between human speech and digital systems, making interactions with technology more seamless, natural, and accessible for everyone.
The ARCHER project organises yearly evaluation campaigns to assess four language technologies: Automatic Speech Recognition (ASR), Machine Translation (MT), Named Entity Recognition (NER), and Optical Character Recognition (OCR). Each campaign covers four languages, two carried over from the previous year and two new ones.
ELDA is in charge of preparing the data used in these evaluations. This includes selecting sources, cleaning and structuring the data, and delivering development and test sets for each technology and language.
The first year includes a “dry run” campaign, launched in early June 2025, to test the processes before the regular yearly evaluations begin. It focuses on English and French for all four technologies.
Dissemination
News from ELRA

LREC 2026, the 15th edition of the Language Resources and Evaluation Conference, will be held in Palau de Congressos de Palma, Palma de Mallorca (Spain), on 11-16 May 2026.
The three-day main conference will take place on 13-15 May 2026, and it will be accompanied by three days of workshops and tutorials to be held in the days immediately before (11-12 May 2026) and after (16 May 2026).
The hybrid conference will bring together researchers and practitioners in natural language processing, computational linguistics, speech and multimodality, with special attention to evaluation and the development of resources that support work in these areas. Following the tradition of LREC, the 15th edition will feature grand challenges and provide ample opportunity for participants to exchange information and ideas through both oral presentations and extensive poster sessions, complemented by an exciting social program
Check the calls on LREC 2026 Website
Assigning DOIs to LREC Papers
The ELRA Board has decided to assign DOIs to all LREC papers published in the proceedings of all LREC editions, starting from LREC-COLING 2024 backwards.
Language Resources and Evaluation Journal
Language Resources and Evaluation is the first publication devoted to the acquisition, creation, annotation, and use of language resources, together with methods for evaluation of resources, technologies, and applications. The Journal is edited by ELRA and published by Springer.
Since January 2025, the following issues have been published.
- Issue 59-4 December 2025
- Issue 59-3 September 2025
- Issue 59-2 June 2025
- Issue 59-1 March 2025
Each of these regular issues include a number of papers in Open Access.
