Corpus Design
The diatopic dimension is traditionally considered the most significant for describing the sociolinguistic variation of contemporary Italian; indeed, traces of regional linguistic features can be found even in the most controlled productions of educated speakers.
To date, linguistic data in the KIParla corpus have been collected in the cities of Bologna and Turin; the sociolinguistic situation of these two survey points is characterized by the co-presence of Italian and dialect, and by the existence of intermediate varieties. Furthermore, although with significant differences, both cities have been and continue to be destinations for internal migration, and for this reason, various regional Italian varieties and Italo-Romance dialects can be found. For this reason, in addition to information regarding the recording location, data on the origins of individual speakers are also accessible.
The speakers involved in the recordings are primarily differentiated by age and educational qualification; both parameters are considered particularly significant for describing the sociolinguistic variation of Italian.
The corpus includes various types of interactions characterized by:
- symmetric/asymmetric relationship between participants;
- presence/absence of a predefined topic;
- presence/absence of rules governing turn-taking
Corpus Construction: Data Collection, Transcription, and Accessibility
All data were recorded with an overt microphone, and all speakers signed an informed consent form (drafted in compliance with current European data protection regulations – see G.D.P.R.) authorizing:
- data collection;
- data storage on hardware located in European countries and/or on cloud services provided by universities;
- the online publication of data for scientific research purposes.
Before being uploaded online, the data (both audio files and transcripts) have been anonymized, and the only directly accessible sensitive data is the speaker's voice itself.
The recordings were transcribed using the ELAN software, which allows for the alignment of the transcription with the audio track.
For the transcriptions, a simplified version of the Jefferson system (see tab. 1) was adopted, which is frequently used in conversation analysis.
| , | Rising intonation |
| . | Falling intonation |
| : | Prolonged sound |
| (.) | Short pause |
| > ciao < | (More) rapid pronunciation |
| <ciao> | (More) slow pronunciation |
| [ciao] | Overlaps between speakers |
| (ciao) | Text of difficult comprehension (transcriber's hypothesis) |
| xxx | Unintelligible text |
| ((laughs)) | Non-verbal behavior |
| = | Prosodically unified units |
Table 1 – Symbols for transcription
The transcribed data, however, can also be searched based solely on the simple orthographic transcription.
Incremental Modularity
Bla
The KIP module
The ParlaTO module
Future prospects
Bla
English Version
An (extended) English version can be downloaded here.