There will be three keynote speakers, in addition to plenary sessions for the Presidential Address and Young Career Award.
Andreas Frey
Goethe University Frankfurt (Germany)
Selecting Empirical Prior Distributions for the Ability Parameters in Large-Scale Educational Assessments
Large-scale assessments (LSAs) such as PISA, PIRLS or TIMSS use complex sampling designs which make it possible to obtain sound estimates for population level statistics and their standard errors. Examples for such population level statistics are country-specific means, percentages of students on proficiency levels, differences of average proficiency between students with and students without migration background, and correlations between proficiency and the socio-economic background of the students. The current multi-step procedure used by all major LSAs is complex and involves IRT scaling, conditioning latent proficiency on background variables, drawing plausible values and aggregating them. Sequential Bayesian adaptive testing is a straight-forward alternative to this multi-step procedure. In Bayesian adaptive testing, background information can be included as prior distribution, item selection and constraint management can be carried out highly efficiently, and population statistics of interest can be derived directly from the posterior distributions. In my presentation, I will summarize the current LSA approach and how Bayesian adaptive testing can optimize and facilitate it. The main focus of the talk will be the question of which prior distribution should best be used to derive unbiased and precise estimates for the population statistics of interest. Results from a comprehensive simulation study will be presented comparing the multi-stage testing design of PISA 2018 with Bayesian adaptive testing implementations with prior distributions of varying complexity, ranging from weakly informative priors to empirical priors incorporating extensive background information. Finally, I will discuss the potential of BAT to extend the reporting capabilities of LSAs and talk about whether applying this approach in operational LSAs is realistic or a playground for psychometricians.
Read more about the keynote here.
Hua-Hua Chang
Purdue University (United States)
Positioning Computerized Adaptive Testing as a Foundation for Personalized Learning
Over the past 70 years, the theory and methods of computerized adaptive testing (CAT) have advanced substantially, and modern technologies have made large‑scale CAT implementation easier and more efficient than ever before. At the same time, the rise of artificial intelligence (AI) in education is reshaping instruction and assessment, exposing the limitations of traditional practices and accelerating the demand for innovative, personalized learning, creating new opportunities for CAT to serve as a core infrastructure for personalized assessment. Yet CAT remains relatively unfamiliar to many researchers working in generative AI (Gen‑AI). This presentation examines the potential of CAT to enhance personalized learning pathways and deliver tailored insights to students. We illustrate how CAT‑based assessments can optimize individualized learning experiences while balancing diagnostic detail with concise, instructionally meaningful feedback. Several large‑scale applications in college STEM instruction will be highlighted to demonstrate how CAT can help us better understand what students know—and how they learn.
Read more about the keynote here.
Maomi Ueno
The University of Electro-Communications (Japan)
High-Stakes Assessments with Process-Integrated IRT and AI: From Empirical Evidence to Computational Design Strategies for Computerized Adaptive Testing
This keynote presents a comprehensive trajectory of high-stakes educational assessment, bridging large-scale empirical success with cutting-edge artificial intelligence and psychometric theories. First, we demonstrate the operational triumphs at the University of Electro-Communications (UEC), Japan, which has successfully administered high-stakes university admissions utilizing computer-based testing (CBT) powered by a newly developed Process-integrated IRT framework. This system evaluates the fine-grained problem-solving processes of applicants in programming, data science, and mathematics. By capturing rich procedural data, such as behavioral logs in programming and step-by-step mathematical formulas via a dedicated expression editor, and analyzing them under the Generalized Partial Credit Model (GPCM), we achieved a two-to-threefold increase in measurement precision (test information volume) even for time-consuming thinking items. Crucially, utilizing this framework for high-stakes university admissions has yielded extraordinary real-world outcomes. Longitudinal institutional data reveals that students admitted through this model achieved a dramatic increase in freshman credit acquisition rates, reaching 100% in informatics and 95% in mathematics, and significantly outperformed general admission cohorts. Furthermore, we introduce predictive learning analytics leveraging Deep-IRT, which successfully forecasts post-enrollment academic performance and credit retention with a high classification accuracy, yielding an Area Under the Curve (AUC) of approximately 0.7.
In the second half of this presentation, we transition from these proven operational successes to forward-looking research designed to overcome the inherent trade-offs of computerized testing. Specifically, we share our latest findings on "Addressing the Accuracy-Exposure Trade-off: Computational Design Strategies for Computerized Adaptive Testing," presenting advanced mathematical frameworks that minimize test length while preventing item over-exposure to ensure secure, large-scale deployment. We also introduce an enhanced Deep-IRT model that explicitly incorporates response-time prediction. Finally, we unveil a groundbreaking approach to Automated Item Generation (AIG) driven by generative AI, capable of autonomously synthesizing novel programming items with pre-specified IRT difficulty parameters and Bloom’ Taxonomy levels. By combining solid empirical data science with generative assessment theory and computational CAT design, this address delineates a sustainable and secure paradigm for educational evaluation in the AI era.
Read more about the keynote here.
Ricardo Primi
Universidade São Francisco (Brazil)
Navigating the Psychometric AI Revolution: Innovations in Scoring and Validity
Artificial Intelligence is driving a paradigm shift in educational and psychological measurement, offering powerful, scalable solutions for automated item generation, computerized adaptive testing, and complex performance scoring. This keynote explores how Large Language Models (LLMs) and natural language processing are redefining traditional psychometric boundaries. Bridging theory and practice, the presentation highlights cutting-edge empirical applications: first, a comparative analysis of AI strategies used for the automated scoring of the PISA 2022 Creative Thinking assessment; and second, the introduction of embeddcv, an R package leveraging text embeddings to evaluate the content validity of psychological scales and map occupations to O*NET categories. Critically examining these use-cases, the session tackles the fundamental issues of validity, reliability, and algorithmic bias, offering a forward-looking roadmap for integrating AI into operational assessments.
Read more about the keynote here.
Wim J. van der Linden
University of Twente (Netherlands)
(Presidential address)
Simplifying large-scale educational assessments through Bayesian adaptive testing
The primary goal of large-scale educational assessment is to monitor the achievements of a population of students through data obtained from blocks of test items administered to samples of students. The standard approach to estimating the achievements has been a plausible values methodology based on a combination of an IRT model for the students responses and a population model relating their abilities to background variables. It is shown how important gains in precision, time, and resources can be realized if the current fixed blocks of items are replaced with adaptive tests running on sequential Bayesian statistics, with a population model serving as initial empirical prior distribution for the ability parameters, and Gibbs sampling of their posterior distributions. The advantages include more realistic updates of the ability parameters right from the beginning of the test, more informative plausible values for the students abilities immediately available at the end of their test, no need to estimate complicated population models conditional on item parameters fixed at point estimates, better student motivation through the option of individual scores reported back to the students, avoidance of time-intensive assembly of large numbers of blocks of test items prior to the assessment, and the option of simultaneous online calibration of new items directly on the operational ability scales in use for the assessment.
Read more about the keynote here.
