RESEARCH ARTICLE Open AccessWhy item response theory should be usedfor longitudinal questionnaire data analysisin medical researchRosalie Gorter 1,2* , Jean-Paul Fox 3 and Jos W. R. Twisk 1,2AbstractBackground: Multi-item questionnaires are important instruments for monitoring health in epidemiologicallongitudinal studies. Mostly sum-scores are used as a summary measure for these multi-item questionnaires. Theobjective of this study was to show the negative impact of using sum-score based longitudinal data analysisinstead of Item Response Theory (IRT)-based plausible values.Methods: In a simulation study (varying the number of items, sample size, and distribution of the outcomes) theparameter estimates resulting from both modeling techniques were compared to the true values. Next, the modelswere applied to an example dataset from the Amsterdam Growth and Health Longitudinal Study (AGHLS).Results: The results show that using sum-scores leads to overestimation of the within person (repeatedmeasurement) variance and underestimation of the between person variance.Conclusions: We recommend using IRT-based plausible value techniques for analyzing repeatedly measured multi-itemquestionnaire data.Keywords: Longitudinal data, Hierarchical model, Item response theory, Questionnaires, Measurement error, Structuralmodel, Plausible values, Multilevel modelBackgroundIn the field of medical epidemiological research, multi-itemquestionnaires are often used to measure the developmentof a subject’s health status over time. The resulting itemobservations are used as measurements of a continuouslatent variable (i.e. a variable that is not directly observable).Examples of latent variables are health related quality of life[1, 2], and depression [3]. A measurement model is re-quired to describe the relation between the observed cat-egorical item responses (for example, Likert items with fouranswering categories: agree/slightly agree/slightly disagree/disagree) and the continuous latent variable.To make statistical inferences about longitudinal mea-surements of the latent variable a statistical model is re-quired that describes the development of the latent variableover time, while addressing the typical correlations betweenmeasurements of one person. The central question is howto measure the latent variable given the response data, andhow to perform the longitudinal data analysis given themeasured variables. In longitudinal designs, the data has anested structure; i.e. repeated measurements are nestedwithin the subjects. Due to the nested structure, the com-mon independence assumptions between measurementsdo not hold and neither linear/logistic regression nor ana-lysis of variance can be used in a straightforward way [4–8].A multilevel model can be used to model the dependencieswhen there are multiple measurements nested within par-ticipants [9]. This multilevel modelling approach will be re-ferred to as structural modeling to explore differences inlongitudinal analyses with sum-scores and IRT-based scoresas estimates for the latent variable.Two fundamental theoretical frameworks can be usedto measure latent variables given the response data. His-torically, there is classical test theory (CTT), where sum-scores are the estimates of the latent variable. The other,theoretically more advanced framework is item response* Correspondence: r.gorter@vumc.nl1 Department of Epidemiology & Biostatistics, VU university medical center,Amsterdam, Netherlands2 EMGO+ institute for health and care research, Amsterdam, NetherlandsFull list of author information is available at the end of the article© 2015 Gorter et al. This is an Open Access article distributed under the terms of the Creative Commons Attribution License(http://creativecommons.org/licenses/by/4.0), which permits unrestricted use, distribution, and reproduction in any medium,provided the original work is properly credited. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated.Gorter et al. BMC Medical Research Methodology (2015) 15:55 DOI 10.1186/s12874-015-0050-x