The KIParla corpus is a new resource for the study of spoken Italian and is the result of a collaboration between the Universities of Bologna and Turin.
The corpus features several innovative characteristics, such as:
- Access to a large amount of metadata about the speakers and the contexts in which the recorded interactions took place;
- The ability to consult the corpus online and access the entire transcription of each conversation;
- Alignment of the transcription with the audio track.
Corpus Design
Geographical differentiation is preeminent in characterizing the sociolinguistic variation of Italian; indeed, regional features can be observed even in the more controlled productions of educated speakers.
Initially, linguistic data for the KIParla corpus were collected in the cities of Bologna and Turin; the sociolinguistic situation in these two survey locations is characterized by the co-presence of Italian and dialect. Furthermore, despite significant differences, both cities have been and continue to be destinations for internal mobility, as well as external migratory flows; consequently, various regional Italian varieties and Italo-Romance dialects can be found there, in addition to recently immigrated languages. For this reason, in addition to information regarding the recording location, data on the geographical origin of individual speakers are also accessible.
With the addition of the KIPasti module, recordings collected across all geographical areas of Italy have been integrated.
The speakers involved in the recordings are primarily differentiated by age, educational qualification, and occupation, which are particularly significant parameters in determining the social placement of individuals.
The corpus contains various types of interaction (such as semi-structured interviews, dinner table conversations, and, in a university context, lectures and exams), differentiated based on situational parameters: symmetric/asymmetric relationship between participants, presence/absence of a predefined topic, presence/absence of turn-taking norms, etc.
Corpus Construction: Data Collection, Transcription, and Accessibility
All data were recorded with an overt microphone, and all speakers signed an informed consent form (drafted in compliance with current European data protection regulations – see G.D.P.R.) authorizing:
- data collection;
- data storage on hardware located in European countries and/or on cloud services provided by universities;
- the online publication of data for scientific research purposes.
Before being uploaded online, the data (both audio files and transcriptions) were anonymized. The only sensitive data, accessible upon registration, is the speaker's voice itself. Sensitive data were replaced in the transcriptions and masked in the audio files.
The recordings were transcribed using the ELAN software, which allows for the alignment of the transcription with the audio track.
For the transcriptions, a simplified version of the Jefferson system (see tab. 1) was adopted, which is frequently used in conversation analysis.
| , | Rising intonation |
| . | Falling intonation |
| : | Prolonged sound |
| (.) | Short pause |
| > ciao < | (More) rapid pronunciation |
| <ciao> | (More) slow pronunciation |
| [ciao] | Overlaps between speakers |
| (ciao) | Text of difficult comprehension (transcriber's hypothesis) |
| xxx | Unintelligible text |
| ((laughs)) | Non-verbal behavior |
| = | Prosodically unified units |
Table 1. Transcription symbols
To make the entire corpus searchable via NoSketch Engine, a Python script has been developed that allows for:
- using metadata both as search filters and as information related to individual recordings;
- performing searches considering both simple orthographic transcription and Jefferson transcription;
- linking each occurrence to the intonational unit in which it is found;
- consulting each module separately.
Incremental Modularity
A fundamental characteristic that makes the KIParla corpus particularly innovative is its incremental modularity, meaning its internal organization into independent modules and the possibility of adding new modules over time.
The modules are distinct spoken Italian corpora that share the same design and a common set of metadata, transcribed from ELAN and made available through NoSketch Engine. Modules can focus on different dimensions of linguistic variation and collect data from various geographical areas. However, the shared data collection and processing procedure ensures a high level of mutual comparability.
The full accessibility of metadata makes the corpus easily expandable, through the addition of further modules focusing on different geographical, socio-cultural, or communicative aspects, and updatable, through the addition of new data for existing modules. The very nature of the KIParla corpus makes it a potential monitor corpus, open to integrations and updates over time.
To date, the KIParla corpus consists of four modules:
The broader the spectrum of collected interactions and the more socio-geographically differentiated the sample of speakers involved, the more representative the corpus will be of the languages and language varieties spoken in Italy.
We envision the KIParla corpus growing in volume over time following two main directions. On the one hand, we aim to collaborate with existing projects to ascertain whether readily available data collected for different purposes can be adapted to form new modules of the KIParla corpus. The sole requirement in such cases is the traceability and accessibility of (at least) a core set of metadata for speakers (gender, age, geographical origin, education level, and profession) and for the interaction (interview, free conversation, etc.). On the other hand, we would like to initiate new data collections in various regions.
Furthermore, in the future, we plan two annotation phases: lemmatization and POS tagging.
English Version
You can find an extended English description here.
Reference:
Mauri, Caterina, Silvia Ballarè, Eugenio Goria, Massimo Cerruti & Francesco Suriano, (2019) “KIParla corpus: a new resource for spoken Italian”. In: Bernardi, Raffaella, Roberto Navigli & Giovanni Semeraro (eds.), Proceedings of the 6th Italian Conference on Computational Linguistics CLiC-it.