ABSTRACT: The first training school organized by the NexusLinguarum COST Action was held on February 8-12, 2021 and was aimed at students, academics, and practitioners wishing to learn the basics of Linguistic Data Science. 
During the training school, the participants were introduced to a wide range of topics: from Semantic Web, RDF and ontologies, to modeling and querying linguistic data with state-of-the-art ontology models and tools. 
The training school was organized under the umbrella of the EUROLAN series of summer schools and was hosted virtually (online) by several institutions: the Romanian Academy, the Research Institute for Artificial Intelligence in Bucharest and the Institute of Computer Science in Ias,i, as well as the “Alexandru Ioan Cuza” University of Ias,i, Romania. 
The training school was attended by 82 participants. 
KEYWORDS: linguistic data science, linked data for linguistics, language data, NexusLinguarum, COST action, EUROLAN, training school. 
NexusLinguarum - European network for Web-centered linguistic data science, COST action CA18029- was launched at the end of October 2019. 
The goal of the NexusLinguarum action is to promote the study of linguistic data science, for which the construction of an ecosystem of multilingual and semantically interoperable linguistic data is required. 
Training schools are one of the means for reaching this goal, and therefore the NexusLin-guarum core team organized the Introduction to Linked Data for Linguistics online training schoolthat took place from February 8 to 12, 2021. 
The training school was aimed at promoting and teaching the basics of linguistic data science and the related technologies to people from the academia and the industry. 
It was organized under the umbrella of the EUROLAN series of summer schools, which was established in 1993 and covers topics that are particularly relevant to the fields of computational linguistics and natural language processing (NLP). 
The goal of this 15th EUROLAN School was to bring together scholars, teachers and students of linguistics, NLP and information technology to discuss the principles and best practices for representing, publishing and linking linguistic data and the issues that constitute the building blocks in the envisioned multilingual and interoperable web-oriented ecosystem. 
The present contribution summarises the organisation, content and results of this training school and is based on Deliverable D1.1 of the Action. 
The training school has been developed for newcomers as well as for those already having basic knowledge in the fields covered. 
The school provided a comprehensive introduction to the methodologies for representing linguistic resources using semantic web technologies, together with the means to extract knowledge from language resources and exploit it using semantic web query languages and reasoning capabilities. 
The topics addressed in the school were the following: 
– Semantic Web and Linked Data(Berners-Lee et al. 2006); 
– Ontologies: RDF (Resource description framework), RDF Schema (Resource Description Framework Schema, variously abbreviated as RDFS, RDF(S), RDF-S, or RDF/S), Web Ontology Language (OWL),etc.); 
– SPARQL query language- a semantic query language for databases able to retrieve and manipulate data stored in the RDF format; 
– Metadata: DCAT (Data Catalog Vocabulary), VOID (RDF Schema vocabulary for expressing metadata about RDF datasets, etc.); 
– RDF transformation and validation; (Cimiano et al. 2020) 
– Linguistic linked data; (Chiarcos et al. 2013) 
– Lemon-OntoLex (McCrae et al. 2017; Declerck, Tiberius, and Wandl-Vogt 2017; Stanković et al. 2018) 
– Linguistic linked data generation; (Cimiano et al. 2020) 
– Corpora and linked data; (Chiarcos 2012) 
– Linguistic annotations; (Fäth et al. 2020) 
– NLP Interchange Format; (Hellmann et al. 2013) 
– Tools and applications of linguistic linked data. (Declerck et al. 2020) 
The first day started with an opening session and a brief introduction to Linguistic Linked Data (LLD), followed by an introduction to Linked Data and RDF dedicated sessions. 
The second day covered topics related to ontologies, including modelling knowledge with ontologies, OWL and SKOS knowledge representation languages, reasoning of knowledge, and a hands-on session using the Protégé ontology editor. 
The third day was dedicated to the topics related to representing and querying lexical data with dedicated sessions on the OntoLex-Lemon model and the SPARQL querying language. 
The fourth day included sessions which gave an overview of other linguistic and metadata vocabularies and the VocBench platform (Stellato et al. 2020) modelling linguistic datasets. 
In the afternoon, an online social event was organized where the participants could remotely see the beauty of the Romanian culture, traditions and nature. 
The fifth day comprised three parallel sessions on different topics: 
(i.) LLD Generation/Transformation and Linking, 
(ii.) Annotations (NIF, Web Annotation) (Hellmann et al. 2013), and 
(iii.) OntoLex extensions: vartrans for representing translations and term variants (based on the lemon translation module, (Gracia et al. 2014)), lexicog – lexicography module (Bosque-Gil, Gracia, and Montiel-Ponsoda 2017), FrAC – frequency, attestation and corpus Information (Chiarcos et al. 2020). 
Finally, the training school ended with a closing session where an ontology of participants, lecturers and organizers was presented, illustrating many of the representation mechanisms explained throughout the week. 
Figure 1. Ontology of the training school. 
Each of the organized sessions was accompanied by a hands-on session and an exercise session. 
During the hands-on session, the lecturers proposed an exercise and offered a step-by-step walk-through for the participants to understand the methodology leading towards the solution. 
They also introduced the basic technology needed. 
Then, during the exercise session, the participants were asked to work on a particular task like the cases presented during the hands-on session, thereby becoming familiar with the technology introduced in a practical setting. 
As these sessions were graded in terms of complexity, starting with the basic notions, and building on to present more specific topics in a detailed fashion on the last day, the participants had a chance to acquire a solid foundation before moving onto more complex sessions. 
The offcial program of the school is available online. 
As a follow up, the JeRTeh Language Resources and Technologies Society set up a local installation of VocBench and, apart from JeRTeh members, it was used by students and teachers of the Intelligent Systems PhD program at the University of Belgrade for the subjects Knowledge repre-sentation and Semantic web. 
The Lemon-OntoLex Frac module was used for representation of the entries from the lexicon used for abusive speech detection with attestations from the Twitter corpus with annotation of abusive spans (Jokić et al. 2021). 
Due to the COVID-19 pandemic and current travel restrictions in Europe and beyond, the training school was held online. 
Following on the almost three decades long tradition of EUROLAN, which is known for academic program excellence and camaraderie among professors and students, a range of virtual activities were carried out in addition to holding online classes, with the aim of providing cultural experiences and discoveries, in addition to closer interaction. 
Attendance was online and free of charge, requiring preregistration. 
All the sessions were hosted using a videoconferencing platform. 
For the hands-on sessions, several breakout (virtual) rooms were made available where the participants could work on the assignment in smaller groups. 
To encourage participants to ask questions and get in touch with each other, the organizers set up a Slack channel, as a collaboration hub where lecturers and participants could clarify any doubts. 
The total number of participants was 82, 52 female and 30 male, including 4 participants from Serbia. 
Various types of materials were generated for the training school, including presentations (slides) and exercises accompanied by code and data examples. 
All the materials were published online and made available for free. 
The training school provided valuable knowledge and trained many computer scientists and linguists on how to work with and benefit from linguistic linked data. 
This was the first training school organized by the NexusLinguarum COST Action as one of a series of training events that are planned to take place. 
It aimed to serve as an introduction to the topic of linguistic data science and build the basis for the audience necessary for attending future training schools on more advanced topics for the duration of the COST Action. 
All the materials created during the training school are publicly available and can be further used by the community. 
During the closing session, the organizers provided participants with a survey form to gather feedback on both organizational and academic aspects of the school. 
The results have shown that the disciplines of the humanities/linguistics/lexicography had a higher representation among participants than computer science, and that the school was well-focused, well-balanced topic-wise and well organized. 
Theory sessions, tutoring, and the opportunities to learn were very highly evaluated. 
On the other hand, due to the virtual mode, there is still room for improvement in practical sessions, social event organization and opportunities to network. 
The knowledge and skills acquired there will improve the development of Serbian linguistic resources and help to publish more resources as linguistic linked data. 
This paper is supported by the COST Action CA18209 - NexusLinguarum “European Network for Web-centred Linguistic Data Science”. 
