General Information
The KIParla corpus currently comprises 4 modules, namely:
- KIP: 668,581 tokens
- ParlaTO: 561,388 tokens
- KIPasti: 482,887 tokens
- ParlaBO: 701,354 tokens
- TOTAL: 2,326,171 tokens*
*the total does not correspond to the sum of the three different modules since KIP and ParlaTO share 7:45
The modules can be consulted independently or via the joint consultation mode. In the latter case, the entire KIParla corpus is accessible.
Overall, the KIParla corpus currently exhibits the structure shown in the figure below.

Metadata
The joint consultation mode provides access to a range of metadata. Unlike the consultation of individual modules, in this mode, the values of various parameters can be displayed in an aggregated fashion.
Conversations:
Interaction type:
- Kitchen table conversation
- Free conversation
- Exam
- Semi-structured interview
- Lesson
- Office hour
Number of participants:
- 1
- 2
- 3
- 4
- 5
- 6
Relationship between participants:
- Symmetric
- Asymmetric
Moderator Presence:
- Yes
- No
Collection year:
- 2017/2018
- 2019
- 2020
- 2021
- 2022
- 2023
- 2024
Collection location:
- AN
- AR
- BO
- BR
- BZ
- CA
- CE
- CZ
- FC
- FG
- LE
- LU
- MI
- MT
- PE
- PG
- RE
- RM
- RN
- TO
- TV
- VE
Speakers:
Occupation:
- N/A (i.e., data not available)
- Artisan
- Merchant
- Unemployed
- Occupations
- Intellectual
- Non-qualified
- Worker
- Pensioner
- Student
- Technical
- Office
Gender:
- Female
- Male
Region of origin:
- Abruzzo
- Basilicata
- Calabria
- Campania
- Emilia Romagna
- Friuli Venezia Giulia
- Lazio
- Liguria
- Lombardy
- Marche
- Molise
- Piedmont
- Apulia
- Sardinia
- Sicily
- Tuscany
- Trentino-South Tyrol
- Umbria
- Aosta Valley
- Veneto
- Foreign
Age:
- 16-20
- 21-25
- 26-30
- 31-35
- 26-40
- 41-45
- 46-50
- 51-55
- 56-60
- 61-65
- 66-70
- 71-75
- 76-80
- 81-85
- Over85
Educational qualification:
- N/A (i.e., data not available)
- Dip_lic (high school diploma)
- Dip_tec_prof (technical or vocational institute diploma)
- Elementary
- Degree
- Degree in progress
- Middle School
- PhD
How to cite the corpus
Mauri, Caterina, Silvia Ballarè, Eugenio Goria, Massimo Cerruti & Francesco Suriano, 2019, “KIParla corpus: a new resource for spoken Italian”. In: Bernardi, Raffaella, Roberto Navigli & Giovanni Semeraro (eds.), Proceedings of the 6th Italian Conference on Computational Linguistics CLiC-it.
Silvia Ballarè, Eugenio Goria and Caterina Mauri, 2022, Spoken Italian and Linguistic Variation. Theory and Practice in the Construction of the KIParla Corpus, Bologna, Pàtron
Coordinators
Caterina Mauri, Silvia Ballarè, Eugenio Goria and Massimo Cerruti
Last update
2024