{"id":10303,"date":"2025-05-28T11:38:12","date_gmt":"2025-05-28T09:38:12","guid":{"rendered":"https:\/\/euroccitaly.it\/?p=10303"},"modified":"2025-05-28T11:40:42","modified_gmt":"2025-05-28T09:40:42","slug":"webinar-language-modeling-for-resource-limited-languages","status":"publish","type":"post","link":"https:\/\/euroccitaly.it\/en\/news\/webinar-language-modeling-for-resource-limited-languages\/","title":{"rendered":"Webinar: language modeling for resource-limited languages"},"content":{"rendered":"<p data-start=\"273\" data-end=\"507\">On June 11, from 10:00 to 11:00, an exclusive webinar will be held featuring speaker Marek Dobe\u0161, focused on <strong data-start=\"382\" data-end=\"430\">language modeling for low-resource languages<\/strong>, organized by the National Competence Centres for HPC in Slovakia and Italy.<\/p>\n<p data-start=\"509\" data-end=\"912\">The rise of large language models (LLMs), such as GPT and LLaMA, has highlighted a significant challenge: the <strong data-start=\"619\" data-end=\"672\">scarcity of data for less widely spoken languages<\/strong>, including Slovak. This limits the quality of AI models for these languages. Our project aims to overcome this barrier through innovative and advanced strategies, supported by the Leonardo supercomputer, one of the most powerful in Europe.<\/p>\n<p data-start=\"914\" data-end=\"956\"><strong data-start=\"914\" data-end=\"956\">Key strategies of the project include:<\/strong><\/p>\n<ul data-start=\"958\" data-end=\"1557\">\n<li data-start=\"958\" data-end=\"1129\">\n<p data-start=\"960\" data-end=\"1129\"><strong data-start=\"960\" data-end=\"1008\">Bilingual Slovak-English dataset generation:<\/strong> Automatic translation assisted by LLaMA 3.3 70B Instruct to create high-quality datasets for training language models.<\/p>\n<\/li>\n<li data-start=\"1130\" data-end=\"1341\">\n<p data-start=\"1132\" data-end=\"1341\"><strong data-start=\"1132\" data-end=\"1190\">Automated summarization of scientific texts in Slovak:<\/strong> Using Gemini Flash Experimental and the PLOS database, summaries of research articles are produced, enhancing specialized terminology in the models.<\/p>\n<\/li>\n<li data-start=\"1342\" data-end=\"1557\">\n<p data-start=\"1344\" data-end=\"1557\"><strong data-start=\"1344\" data-end=\"1387\">Cultural context enrichment for Slovak:<\/strong> Development of specific datasets to improve the understanding of Slovak cultural and contextual topics, which are still underrepresented by existing models like ChatGPT.<\/p>\n<\/li>\n<\/ul>\n<p data-start=\"1559\" data-end=\"1795\">The project leverages the computing power of the Slovak national supercomputer Devana and the Leonardo supercomputer in Italy to train next-generation language models, thereby improving accuracy and cultural relevance of LLMs in Slovak.<\/p>\n<p data-start=\"1797\" data-end=\"2108\">Although the focus is on Slovak, the methods developed are applicable to many other low-resource languages worldwide. The project invites international collaborators to join in promoting <strong data-start=\"1984\" data-end=\"2038\">European cooperation in high-performance computing<\/strong> and fostering a more inclusive, multilingual artificial intelligence.<\/p>\n<p data-start=\"2110\" data-end=\"2300\"><strong data-start=\"2110\" data-end=\"2130\">Join the webinar<\/strong> to discover how the combination of supercomputing and linguistic innovation can open new frontiers for the development of language models for underrepresented languages!<\/p>\n<p><strong><a href=\"https:\/\/forms.office.com\/e\/UECHKV1gA3\" target=\"_blank\" rel=\"noopener\">Link to sign up<\/a><\/strong><\/p>\n","protected":false},"excerpt":{"rendered":"<p>On June 11, from 10:00 to 11:00, an exclusive webinar will be held featuring speaker Marek Dobe\u0161, focused on language modeling for low-resource languages, organized by the National Competence Centres for HPC in Slovakia and Italy. The rise of large language models (LLMs), such as GPT and LLaMA, has highlighted a significant challenge: the scarcity&#8230;<\/p>\n","protected":false},"author":2,"featured_media":10301,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_members_access_role":[],"_members_access_error":""},"categories":[18,14,15],"tags":[],"class_list":["post-10303","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-networking-en","category-news","category-training"],"_links":{"self":[{"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/posts\/10303","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/comments?post=10303"}],"version-history":[{"count":1,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/posts\/10303\/revisions"}],"predecessor-version":[{"id":10304,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/posts\/10303\/revisions\/10304"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/media\/10301"}],"wp:attachment":[{"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/media?parent=10303"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/categories?post=10303"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/euroccitaly.it\/en\/wp-json\/wp\/v2\/tags?post=10303"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}