Multilingual corpus in HEALTH (COVID-19) domain part_1a (v.1.05) in TSV/MOSES-like format. 
This dataset has been generated out of public content available through several websites of national agencies (https://www.ecdc.europa.eu/en/COVID-19/national-sources) and selected broadact websites like (Global Voices, Voxeurop, voltairenet, etc.)
The dataset contains 327 X-Y TSV/MOSES-like (pairs of) files, where X and Y belong to the set {CEF language plus IS and NO} (3044961 TUs in total). Acquisition of data (from multi/bi-lingual websites), normalization, cleaning, deduplication and identification of parallel documents have been done by ILSP-FC tool. Multilingual embeddings (LASER) were used for alignment of segments. Merging/filtering of segment pairs has also been applied.
DSI Relevance: eHealth
People who looked at this resource also viewed the following:
- Compilation of German-Estonian parallel corpora resources used for training of NTEU Machine Translation engines.
- Compilation of Irish-Polish parallel corpora resources used for training of NTEU Machine Translation engines. Tier 3.
- Compilation of Czech-Estonian parallel corpora resources used for training of NTEU Machine Translation engines.
- Bilinguis Free Books v.1.04. Multilingual (CS, DE, EN, ES, FI, FR, IT, NL, PL, PT) corpus from the http://bilinguis.com/ website.