ABSTRACT: 
In this paper, we discuss the lessons learned through the lifecycle of a di-alectal electronic lexicon. 
Our approach is in-novative because our lexicon is designed and built as a multi-dialectal (trilingual) dictio-nary (three dialects vs. one target language) instead of three monolingual dialectal dictionaries. 
Our system oﬀers features that could not be possible with three monolingual dialectal dictionaries. 
Moreover, during the system’s lifecycle we have got very specific demands for improvements (new requirements) that users were not able to express during the analysis phase. 
The lessons learned and the solutions invented for the system’s ultimation (to re-spect the new requirements) can be helpful for other research or project with similar pur-poses.  
KEYWORDS: Computational Dialectology, Dialectal Lexicography, Electronic Dictionaries, Lexical Resources, Modern Greek Dialects, Asia Minor Greek.  
In a previous work (Karanikolas et. al., 2013), we have presented the design and implementation of a multimedia electronic dictionary of three Greek dialects in Asia Minor (Pontic, Cappadocian, Aivaliot). 
We had pre-sented the linguistic and lexicographic approach adopted, as well as the principles for designing the macro/microstructure of the dictionary. 
We also had presented the conceptual model of the tri–dialectal dictionary and the equivalent relational schema. 
According to the above analysis a system has been implemented that hosts lemmas and relevant lexicographic information from three Asia Minor Greek dialects.  
However, during the lifecycle of the system, and because of the highly– qualified users, we have got very specific demands for improvements and we have caught the ultimate goal (an excellent system). 
In this paper, we report the lessons learned through this system’s lifecycle and we present the improved design and the extended facilities of the 3–dialectal dictionary. 
We claim that our system can be used for other Greek dialects and that our extended design can be the base for multi–dialectal dictionaries for other languages. 
Our ultimate system can be used for multi–dialectal dictionaries of other languages so long as other virtual keyboards can be appended to it.  
The paper is organized as follows. 
Section 2 presents the motivation for building the 3–dialectal lexicon and relevant work is presented in section 3. 
Sections 4 and 5 describe the initial requirements and the design according to the requirements set. 
Section 6 gives some details from the first implemented version of the system. 
Section 7 presents the demands for improvements and the relevant implementations. 
The result of the improvements is an excellent system and the design of this system is the topic of section 8 while conclusions are drawn in section 9.  
Pontic, Cappadocian and Aivaliot are three Greek dialects in Asia Minor which are not suﬃciently documented and they are on the way to extinction. 
Until now, little interest has been shown in the dialects in question. 
The most interesting exception is the Papadopoulos’ historical dictionary of Pontic (Papadopoulos, 1958). 
We can also find mentions to Cappadocian in some other works (Thomason, 2001; Thomason and Kaufman, 1988). 
There are also some glossaries for the Asia Minor Greek dialects containing words and idiomatic phrases accompanied by their meaning in Standard Modern Greek. 
However, in most of these glossaries, lemmas are stored in a very unsystematic way and crucial information, such as pronunciation or usages, is missing. 
Moreover, some verbs are listed in their past tense form while others appear in the present tense. 
Therefore, a sound linguistic analysis of Asia Minor Greek dialects is indispensable and gives insights as for the nature and mechanism of language change within the domain of dialectal variation.  
This and other relevant social speculations (syllogisms) motivated us for the initiation of the AMiGre project, within the framework of THALIS pro- gram. 
The project acronym (AMiGre) comes from the project’s title: “Pon-tus, Cappadocia, Aivali: in search of Asia Minor Greek”. 
One of the de-liverables of the AMiGre project was the design and implementation of a multimedia tri–dialectal dictionary for three Greek dialects in Asia Minor (Pontic, Cappadocian, Aivaliot), which we discuss in this paper.  
Dialectal dictionaries are usually treated as monolingual synchronic dic-tionaries. 
In our case (AMiGre), instead of creating three monolingual dialec-tal dictionaries, we have decided to treat and design a trilingual dictionary (three Asia Minor dialects vs. Standard Modern Greek). 
This is the most interesting technical motivation. It is also an interesting innovation because, it is permitting cross–reference links from lemma to lemma (of the same or different dialect) and equivalence links between meanings of lemmas from diﬀerent dialects. 
These could not happen with three monolingual dialectal dictionaries.  
Electronic lexicography for Modern Greek was not concerned with the creation of dialectal dictionaries until very recently. 
The online dictionaries developed at the Portal for the Greek Language (Online, 2016) comprise the computerized versions of Georgacas’ Greek–English Dictionary, Triandafyl-lides’ Dictionary of Standard Modern Greek and Anastasiadi–Symeonidi’s Reverse Dictionary. 
In addition, the Portal provides access to the comput-erised version of Kriaras’ Concise Dictionary of Medieval Vulgar Greek Liter-ature. 
The Institute for Language and Speech Processing has developed on-line bilingual dictionaries (Greek–English, Greek–German, Greek–Russian, Greek–Turkish, and Greek–Arabic). 
The dictionaries are under continuous development and enhancement and they are available from (ILSP, 2016). 
In addition, NLP tools for supporting lexicographic applications have been developed. 
Indicatively, in (Tsalidis et. al., 2010) infrastructure tools which are used for encoding morphological, syntactic and semantic information are reported as well as proofing tools such as a spelling checker, a hyphenator etc. 
As far as Greek dialects are concerned, the only computerized dictionary to our knowledge is the online lexical database of Cypriot Greek (Themis-tocleous, 2012). 
The online dictionary environment provides an enhanced searching mechanism as well as text to speech features for the pronunciation of Cypriot Greek words.  
Dialectal dictionaries are usually treated as monolingual synchronic dic-tionaries (B´ejoint, 2000; Geeraerts, 1989), due to limits in macrostructure (overall organizational scheme of lemmas) (Landau, 2001; Zgusta, 1971). 
Given that our purpose was to design and build an online dictionary, its macrostructure will not be restricted by physical constraints (limitations ex-isting for print dictionaries), and could oﬀer (virtually) “multiple macrostruc-tures” mirroring the various searching options that we could build (Burke, 2003). 
Therefore, since there were no limits in macrostructure, we have de-cided to design and build a trilingual dictionary (Three Asia Minor dialects vs. Standard Modern Greek) (Xydopoulos and Ralli, 2012), instead of three monolingual dialectal dictionaries. 
The dictionary is named TDGDAM (Tri– Dialectal Greek Dictionary of Asia Minor) and it aims to be a linguistically– sound tri–dialectal dictionary in electronic form. 
One basic requirement of TDGDAM was that users should have access to a graphic (form based) rep-resentation of each lemma permitting them to handle pronunciation, mean-ing, usages and relations with other lemmas. 
The representation should be editable and for this to be possible, conventionally–adopted character sets should be used. 
Among other things, each lemma should contain the dialec-tal area and the source from which the lemma has been extracted. 
This type of dictionary constitutes an innovation not only for the Greek language and its dialects, but also for the international standards, as will be explained below.  
Regarding its geographic and time scope, TDGDAM was designed to be a local/ microareal dialectal dictionary of non–synchronic nature that should include entries from diﬀerent areas and time periods (Penhallurik, 2009). 
As it was decided from the beginning, the lemmas of TDGDAM should be drawn (directly or indirectly) from oral speech and written material of the particular dialectal varieties (Keymeulen, 2010).  
Regarding TDGDAM’s microstructure, our aim is to include formal in-formation about pronunciation (phonetic form), grammar (categorial and morphological information), origin (etymology), meaning (synonymic and/or descriptive definitions), usage (thematic and register labels) and to provide linked multimedia resources (internal or external to TDGDAM) to enrich the semantics and pragmatics of lemmas (Barbato and Varvaro, 2004; Rys and Keymeulen, 2009; Xydopoulos and Ralli, 2012). 
To avoid different and arbitrary spelling codes for the same dialect (Durkin, 2010; Xydopoulos, 2012), headwords do not appear in a “semi–phonetic” transcription but in (capitalized) orthographic form. 
In particular, the capitalized orthographic form departs from the spelling form in the standard dialect; it does not prescribe spelling rules in the dialect and allows for any alternative ortho-graphic forms to appear in microstructure (Markus and Heuberger, 2007; Xydopoulos, 2012). 
Finally, authentic examples of use were considered as es-sential constituent information in entries which will appear in non–standard spelling, reflecting pronunciation as closely as possible with the use of dia-critics, but avoiding a “semi–phonetic” transcription (Rys and Keymeulen, 2009).  
Regarding the abilities for cross linking between items of the TDG-DAM, we have defined 3 necessities: 
Cross–reference to other entries, re-lated either through derivational processes or through semantic relations; Equivalence links between meanings of lemmas from diﬀerent dialects; Syn-onymic/Antonymic relations.  
The following 3 figures (figures 1, 2 and 3) present draft structural de-pictions of an equivalent number of lemmas that TDGDAM should contain. 
Based on these and other similar draft structural depictions we designed the TDGDAM system.  
The terms synonymy (Synonym, 2016) and antonymy (Opposite, 2016) used previously and the terms homonymy (Homonym, 2016) and polysemy (Polysemy, 2016) that will be used later are very well defined. 
Their defini-tions are available on the internet.  
Based on the analysis presented in the requirements section and the draft structural depictions (see figures 1, 2 and 3) the following structure of lemmas is the result: 
– Headword, dialect (dialectal region), morphological information/process and etymology are primary information with single values that together define and are dependent on the lemma. 
– Each lemma can have many different realizations and each one of them is characterized by a slightly different phonetic realization dependent on the micro–dialectal region it originates from (the specific area within the wider dialectal region where the lemma’s realization occurs). 
– Each lemma can possibly have diﬀerent meanings (i.e. polysemy), or be homonymous with other, semantically distinct, lemmas. 
– For each meaning, diﬀerent usage examples are essential.  
Regarding the relations (between lemmas and meanings of lemmas) we concluded the following: 
– Cross reference (“See also”) links can be available for connecting lemmas that are semantically / pragmatically / morphologically / etymologically related to each other. 
– Synonyms and Antonyms are two semantic relations that apply between lemmas. 
Both relations relate a lemma meaning with a lemma (the refer-enced one). 
Synonym and Antonym links are restricted between a lemma meaning and a lemma from the same dialect. 
– There are meanings of diﬀerent lemmas from diﬀerent dialects that share the same definition. 
This relation is labeled “Other Dialect”. 
In contrast with the rest of the relations, “Other Dialect” is a symmetrical relation.  
The overall idea (lemma structure and relations) is strictly defined as it is depicted with the Entity Relation Diagram of figure 4.  
The following four data dictionaries (tables 1, 2, 3, 4) explain the four sections (sub–schemas) of the overall ERD. 
One possible implementation of the conceptual model (ERD) using a rela-tional database is depicted in the relational schema of figure 5 which contains thirteen tables. 
However, only seven tables are important. 
The other six ta-bles are lookup tables (listing the set of available values existing) related to some fields of the important tables. 
The important tables are highlighted (in figure 5) with thicker border and larger font in their title. 
Four out of seven important tables are the relational equivalents of the main conceptual entity (“Lemma”), the weak entity (“meaning”) and the two multiple-valued com- posite attributes (“Realization Types” and “Usage Examples”). 
The remain-ing three important tables are the relational equivalents of the conceptual relations (“See Also”, “Thesaurus” and “Other Dialect”).  
Only the table MeaningSets (the implementation of the conceptual rela-tion “Other Dialect”) needs more explanation. 
This relation is symmetrical by nature, i.e. whenever a meaning of a certain lemma from one dialect is declared as being the equivalent of the meaning of another lemma from a diﬀerent dialect, then the reverse is implied. 
It is the structure of table Mean-ingSets and the application’s logic that assures this symmetry. 
The other two relations (“See Also” and “Thesaurus”) are not symmetrical by nature. 
This is reflected in the relational schema (and the application logic). 
Consequently, the user must define the relation in both directions, in case an instance of them (the “See Also” or the “Thesaurus” relation) is symmetrical.  
The International Phonetic Alphabet (IPA) which is used in Table 2 – sub–schema for “Realization Types” – is an alphabetic system of phonetic notation based primarily on the Latin alphabet. 
It was devised by the Inter-national Phonetic Association as a standardized representation of the sounds of spoken language. 
The IPA is used by lexicographers, foreign language stu-dents and teachers, linguists, speech–language pathologists, singers, actors, constructed language creators, and translators. 
Figures 6 and 7 present the most useful IPA charts.  
Another approach to phonetic notation is SAMPA (Speech Assessment Methods Phonetic Alphabet) and it is a machine–readable phonetic alpha-bet. 
It was originally developed under the ESPRIT project 1541, SAM (Speech Assessment Methods) in 1987–89. 
It applied first to Danish, Dutch, English, French, German, and Italian (1989). Later, it applied to Norwegian and Swedish (1992). 
Subsequently it applied to Greek, Portuguese, and Spanish (1993). It has now been extended to Bulgarian, Estonian, Hungar-ian, Polish, and Romanian (1996).  
The GUI version of the system is based on two forms: “main form” and “meaning form”. 
Figure 8 presents the main form for the lemma “ALLOUGURISTRA”. 
The main form is divided into 3 sections. 
The up-per section provides information on the headword, etymology, morphological process and dialect. 
The middle section is a two–card panel. 
The first card in the panel is used for displaying and editing realizations, while the second one is used for providing the meanings list of lemmas. 
The lower section of form is a panel for hosting the “see also” reference list. 
A more detailed de-scription of main form’s middle section is provided in figure 9 which depicts the second card (meanings list) for the same lemma.  
The “meaning form” of a lemma is invoked by an action button once the user selects an item from the “meanings list” of “main form”. 
Fig-ure 10 depicts the “meaning form” presenting certain meanings of the lemma “ALLOUGURISTRA”. 
The meaning form is divided into 3 sections. 
The up-per section provides the definition of the meaning, optionally a picture and the usage label. 
The middle section of the form is a panel for hosting the “us-age examples” list. 
The lower section of the form is a two-card panel. 
The first card in the panel is used for displaying and editing synonymic/antonymic re-lations (thesaurus), while the second one is used for handling the equivalents in other dialects.  
According to Analysis (Requirements) and Design sections, the imple-mented system provides the user with the following character sets for editing the relevant fields: 
– Etymology 
Greek Polytonal 
Loan characters from other alphabets in case of loan words (e.g. characters from the Turkish alphabet) 
– Phonetic type IPA 
– Spelling (Phonetic Orthography) 
Modern Greek 
Accents 
Hyphen, parentheses, apostrophe.  
1. As it is well known, users express most of their arguments during the final stage of lifecycle of the initial development of the system (acceptance, installation, deployment) and during the maintenance stage. 
There was no exception to this rule in the electronic lexicon. 
In our case users explained that the origin (etymology) attribute of a lemma, should be denoted with respect to the sources and the conventions of the originat-ing language (from where the lemma comes). 
Therefore, in the case of a multi–dialectal dictionary, the etymology attribute can contain words from any of the languages that have influenced the dialect. 
Consequently, we came up with a solution which was to provide (in a latter develop-ment) visual keyboards for any aﬀecting language. 
In the case of the 3 dialects of AMiGre the influencing alphabets are Greek, Ancient Greek and Turkish. 
Figure 11 presents a 3–card virtual keyboard with cards for lowercase ancient Greek, uppercase Ancient Greek and Turkish. 
These, together with Modern Greek (provided by the physical keyboard), per-mitted users to enter the required etymology of each lemma without any restriction.  
2. Regarding the phonetic type attribute and the intonation of vowels there are two options available. 
The first option is to place an accent before stressed syllable. 
The second option is to use the stressed version of the vowel. 
Both options are semantically equivalent. 
However, the first option is easier to use and implement in a system because there is no need to support both versions (with and without accent) of the IPA symbols used for the vowels. 
Thus, in the initial development, our sys-tem permitted only the vertical accent before the syllable for denoting the intonated vowel. 
Therefore, in figure 8 the phonetic type is denoted with “aluji’ristra”. 
After the initial development and during the mainte-nance stage a need to support accented versions of vowels came up. 
To cope with this demand, we have extended the system with virtual key-boards for easily inserting any IPA symbol, with and without accents. 
Figures 12, 13, 14 and 15 present the virtual keyboards used for the phonetic type attribute.  
The combinations of IPA symbols (fig. 12) with some of the Diacritic symbols (fig. 13) produce the accented IPA symbols. 
Figure 16 present another lemma that has phonetic type with accented IPA symbols.  
3. According to the linguistic analysis, the “spelling” attribute represents a non–standard graphic representation of pronunciation according to the orthographic rules of Modern Greek (target language), combined with diacritics to annotate any phonological alternations. 
Therefore, we concluded that the domain of values for the spelling attribute can be strings containing the letters of the Greek alphabet and definite diacritic symbols (accent, hyphens, parentheses and apostrophes). 
This conclu-sion was followed for the first implemented version of the system. 
Since the system got into the production stage, we became recipients of very specific demands for improvements in the spelling attribute. 
The users pointed out that some words, since they originate from other languages, have vowels that do not exist in the standard spelling system of the target language (Modern Greek). 
Since the studied dialects have a prominent number of loan words originated from other languages, this remark could not be ignored. 
So, we had to provide some way for the users of system to be able to understand how to pronounce the dialectal words, without engaging them to read phonetic (IPA) symbols. 
The simplest way was to allow the insertion of the grapheme representations (letters) used in the originating language of the loan words for the representation of vowels that do not exist in the target language (Modern Greek). 
The outcome was a small number of vowels with grapheme representations having umlauts (u, e and , with umlaut). 
These letters were included in one virtual keyboard added for editing the spelling (phonetic transcription) attribute. 
Figure 17 represents this virtual keyboard.  
4. As we have already described, each lemma can have many diﬀerent real-izations and each one of them is characterized by a slightly diﬀerent pho-netic realization dependent on the micro–dialectal region it originates.  
This principle drove our design and we built a system where each realiza-tion is characterized by one micro–dialectal area (sub–area of the wider dialectal region). 
However, in the production stage of the system, users pointed out that a single realization can exist in more than one micro– dialectal regions (of the dialectal region). 
To comply with this lately defined requirement we modified the data schema design and the ap-plication. 
The “Microdialectal region” attribute was changed to become multi–valued and the corresponding GUI item changed to hold a list of values (the domain of each value is the set of micro–dialectal regions existing for the dialect of lemma). 
Figure 18 presents the realizations of the Pontic lemma “OMMATOTZATZI” where the third realization (third line) has 3 micro–dialectal regions (TrapezoÔnta, Qald—a, S nta).  
5. Another feature of the system which was not defined in the analysis phase but emerged during the development phase was the content of a “see also” reference list item for a destination lemma having more than one meaning and/or having more than one realization type. 
The solution we decided to follow was to represent the definition of each meaning for the referenced lemma and the spelling of each realization for the referenced lemma in the “see also” list item. 
The lower section of figure 16 represents the empty reference list of a lemma but we can see the plural number used in titles of the relevant columns (Spellings and Meanings).  
6. Another worth–mentioning feature of the system is that Syn-onymic/Antonymic links can refer to another lemma of the same dialect which may or may not be present in the system. 
Figure 19 represents a meaning of lemma “LIWSTRA“ with its “thesaurus” (table providing the Synonymy/Antonymy links). 
In this figure, we can see 3 links (refer-ring lemmas SORTA, TAKIOU and ALLOUGURISTRA) but only the third one is present in the system. 
A comparison of figures 10 and 19 denotes that in the thesaurus table of the newer version (lower section of figure 19) we have replaced Spelling with Etymology. 
Note that we have also added a new column (Source type) in the usage examples table (middle section of figure 19).  
7. So far, we haven’t seen any example of equivalent in another dialect. 
Fig-ure 20 is the main form of the Cappadocian lemma “ANTET“ which has a single meaning. 
In figure 21 we provide the lower section of the meaning form displaying the “equivalent in other dialects” card. 
As depicted in fig-ure 21 the only available meaning of the Cappadocian lemma “ANTET“ has an equivalent meaning in the lemma Aivaliot lemma “ANTETI“.  
8. During the production stage of the system’s lifecycle we have noticed that users diverged from the regulations for writing the usage examples of lemmas. 
Usually users exploited the copy/paste feature of the oper-ating system in order to enter characters not provided directly by the applications for the “usage example” attribute. 
For example, the value “AclageÔw to m lon” was entered in the usage example attribute of the lemma “ASLAEUW". 
This value has Greek characters that are accord-ing to the regulations but also contains a Turkish character (the second character in the string value). 
The phrase “AclageÔw to m lon” as value in the usage example, together with the value “tourk. 
A¸Slamak” (i.e. “from Turkish a¸Slamak”) in the etymology attribute of the same lemma can be an indication of how native speakers of the dialect could possibly write the dialectal word in their documents ("AclageÔw"). 
However, this indication is hidden inside one of the meanings of a lemma. 
We suppose that it could be better to provide another attribute (named “indicative writing”) in each realization of the lemma. 
In this way, the “indicative writing” would be directly available in the main form of lemmas and moreover it would be diﬀerentiated in each realization (micro–dialectal regions). 
This is the only feature that is not implemented in the sys-tem’s ultimation because it is denoted very late, but we consider it very valuable for next multi–dialectal lexicons.  
The data schema (ERD) for supporting the ultimate system is given in figure 22.  
TDGDAM’s projected macrostructure includes ca. 2,500 entries from each of the three dialects of Asia Minor Greek (a total of ca.7,500 entries). 
These entries are drawn from collected vocabulary solely from the three di-alects concerned and exclude all vocabulary found in Standard Greek (unless diﬀerently used). 
Their listing is based on alphabetical, and not onomasio-logical, organization, accessed via dynamic searching options (Xydopoulos, 2012).  
This research is co–financed by the European Union (European Social Fund – ESF) and Greek national funds through the Operational Program “Education and Life–long Learning” of the National Strategic Reference framework (NSRF) – Research Funding Program: “THALIS. Investing in knowledge society” through the European Social Fund. 
We thank Angela Ralli, Professor of Linguistics at the Department of Philology of the Uni-versity of Patras, who is the Coordinator of the whole AMiGre project. 
We also thank George J. Xydopoulos, Associate Professor of Linguistics at the Department of Philology of the University of Patras, who set the linguistic requirements for the dialectal electronic lexicon. 
