Schedule At A Glance
Monday, 28 September
- Morning: Pre-Conference Workshops*
- Afternoon: Welcome Plenary and Paper Sessions
- Evening: Welcome Reception and Poster Session
| DAY 1 - 28/9 (Monday) | ||
|---|---|---|
| Time | Activity | Venue |
| 08:00 - 10:00 | Workshop 1 | Room 1110 |
| Workshop 2 | Auditorium 10th floor | |
| Workshop 3 | Room 1150 | |
| 10:00 - 10:30 | COFFEE BREAK | Room 1120 |
| 10:30 - 12:00 | Workshop 1 | Room 1110 |
| Workshop 2 | Auditorium 10th floor | |
| Workshop 3 | Room 1150 | |
| 12:00 - 14:00 | LUNCH | Free |
| 14:00 - 14:30 | Opening ceremony | Theater |
| 14:30 - 14:45 | Introduction of board members and Welcome to incoming president Chun Wang | Theater |
| 14:45 - 15:45 |
Keynote 1 - Wim van der Linden (Presidential address) Simplifying large-scale educational assessments through Bayesian adaptive testing Chair: Chun Wang |
Theater |
| 15:45 - 16:15 | COFFEE BREAK | Foyer Theater |
| 16:15 - 17:15 |
Keynote 2 - Andreas Frey Selecting Empirical Prior Distributions for the Ability Parameters in Large-Scale Educational Assessments Chair: Jonathan Templin |
Theater |
| 17:15 - 18:00 |
Awards and Travel Grants Chair: Nathan Thompson |
Theater |
| 18:00 - 19:00 | Poster session (9 works) | Foyer Theater |
| 19:00 | Welcome Reception | Foyer Theater |
Tuesday, 29 September
- Morning: Plenary and Paper Sessions
- Afternoon: Plenary and Paper Sessions
- Evening: Conference Dinner — Concert, Tour and Dinner at Catedral da Sé*
| DAY 2 - 29/9 (Tuesday) | ||
|---|---|---|
| Time | Activity | Venue |
| 08:00 - 9:40 | Symposia: EduCAT - chair: Dalton F. Andrade | Theater |
| Adaptive Assessment in Practice: Applications of CAT at SESI and in the ANPAD Test (Franciele Sena) | ||
| Item Calibration with Avatars: Innovation and Equity in Assessment (Dalton F. De Andrade) | ||
| Individual presentations: Applications of CAT (Education and Language) – chair: Luis Valdivieso | Room 1150 | |
| Development and Evaluation of an IRT-Based CAT System for General English Proficiency of College Students (Dong Seo) | ||
| Growth Norms for Computer Adaptive Tests: Evaluating Three Conditional Growth Percentile Methods (Luciana Cancado) | ||
| Enabling Computerized Adaptive Testing in large-scale educational assessments under limited connectivity conditions (Gustavo Silva) | ||
| Development of a Computerised Adaptive Test for Assessing Upper-Basic School Students' Mathematics-Ability in Nigeria: A Post-hoc Simulation Testing (ISAAC OLAWALE IFINJU) | ||
| Designing multi-stage CAT assessments of academic English language proficiency to give examinees the greatest opportunity to demonstrate that they meet performance level thresholds. (Stephen Walker) | ||
| Individual presentations: Other AI Topics in Psychometrics – chair: Joaquim Soares Neto | Room 1110 | |
| Using Semantic Embeddings to Examine Content Overlap Between the IDCP-2 and PID-5 (Pedro Godoy dos Santos) | ||
| AI-Based Item Parameter Estimation as an Alternative to Pilot Testing in CAT: A Simulation Study (Zeus Bellido) | ||
| Adaptive Testing Meets Contracting: Rethinking Success in the Age of AI (Xiangen Hu) | ||
| AI-Assisted Cognitive Pretesting for Personality Assessment Items: Comparing Profile-Conditioned Simulated Responses With Observed Human Data (Diego Forteza) | ||
| Use of AI to Assess Metaphor Generation: Convergent And External Validity (Beatriz Chrispim) | ||
| 9:40 - 10:00 | COFFEE BREAK | Foyer Theater |
| 10:00 - 11:00 |
Keynote 3: Hua-hua Chang – chair: Wim van der Linden Positioning Computerized Adaptive Testing as a Foundation for Personalized Learning |
Theater |
| 11:00 - 12:00 |
Keynote 4: Maomi Ueno – chair: Wim van der Linden High-Stakes Assessments with Process-Integrated IRT and AI: From Empirical Evidence to Computational Design Strategies for Computerized Adaptive Testing |
Theater |
| 12:00 - 14:00 | LUNCH | Free |
| 13:00 - 14:00 | FIESP Cultural Center guided tour | |
| 14:00 - 15:40 | Individual presentations: Other AI Topics in Psychometrics – chair: Richard Gershon | Room 1150 |
| Psychometrics Meets LLM: Diagnostic Assessment Copilot Design and Deployment (Chun Wang) | ||
| Automated Evaluation of Teacher Portfolios Using LLM and Structured Educational Evidence Analysis (Gustavo Silva) | ||
| Methodology for the Empirical Generation and Validation of Performance Level Descriptors Using Generative AI (Daniela Jimenez) | ||
| Symposia: New IRT models to modeling different responses – chair: Jorge Bazan | Auditorium 10th floor | |
| Revisiting the Skew-Normal IRT family: Bayesian estimation and application (Cristian Bayes) | ||
| A robust GNBk count item response model (Luis Valdivieso) | ||
| Bayesian Estimation for a class of ability-based guessing IRT models (Alex de la Cruz Huayanay) | ||
| Symposia: Frontiers of Bayesian Adaptive Testing – chair: Jonathan Templin | Theater | |
| Prior Specification for Ability Estimation in Bayesian Adaptive Testing (Jonathan Templin) | ||
| Spectral Efficiency in Multidimensional Bayesian Adaptive Testing: Comparing Fully Bayesian, Sequential Two-Level, and Globally Optimized Two-Level Frameworks (Seung Choi) | ||
| Bayesian Adaptive Testing for the Answer-Until-Correct-Response Format (Andreas Frey) | ||
| Bayesian Optimal Large-Scale CAT Using Constraint-Governed Augmentation (Richard Patz) | ||
| 15:40 - 16:10 | COFFEE BREAK | Foyer Theater |
| 16:30 - 17:30 | FIESP Cultural Center guided tour | |
| 16:30 - 22:00 | Departure for the Cathedral and Conference Dinner | Free |
Wednesday, 30 September
- Morning: Plenary and Paper Sessions
- Afternoon: Closing Plenary and Paper Sessions
- Evening: Tour in São Paulo*
| DAY 3 - 30/9 (Wednesday) | ||
|---|---|---|
| Time | Activity | Venue |
| 08:00 - 09:00 |
Early Career Award: Dr Hyeon-Ah Kang – chair: Nathan Thompson Beyond Information: Balancing Measurement Precision, Testing Efficiency, and Item Security in Technology-Enhanced Adaptive Testing |
Theater |
| 09:15 - 10:30 | Individual presentations: Cognitive diagnostic and IRT models — chair: Hua-hua Chang | Room 1110 |
| A Cognitive Diagnosis Model for Latent Classification of Bounded Continuous Variables (Eduardo Schneider Bueno de Oliveira) | ||
| Black-Box Variational Inference with REINFORCE for the DINA/DINO Model (Flavio Barros) | ||
| Symposia: Continuous Item Banking for a High-Stakes Certification Without Pretesting: An Operational Decomposition — chair: Andrea Burgos de A. Mangabeira | Theater | |
| Interactive Branching-Scenario Items in High-Stakes Financial Certification: Bundle Scoring Without Venue Dependence (Andrea Burgos de Azevedo Mangabeira) | ||
| Closing the Loop on Item Demand: A Decomposable Formalism for Continuous Item Banking (Marco Pepe) | ||
| Continuous Calibration of an Operational Item Bank Without Pretesting (Erica Ruiz) | ||
| Individual Presentations: Item generation, item banking, form assembly, and others — chair: Alexandre Jaloto | Room 1150 | |
| Form Assembly: An AI Agent Workflow for Building Parallel Spring Forms for a Diagnostic Classification Assessment (Ann Hu) | ||
| Accelerating Item Piloting with NLP-Based Features and Bayesian CAT (Steven Nydick)¹ | ||
| LLM and Rule-Based Automatic Item Generation: A Two-Study Investigation in Large-Scale Literacy Assessment (Lucas Larcher) | ||
| 10:00 - 10:30 | COFFEE BREAK | Room 1120 / Foyer Theater |
| 10:45 - 12:00 | Symposia: The PROMIS, NIH Toolbox, and Mobile Toolbox Measurement Systems: Architecture, Psychometric Foundation, and the Challenges of Adaptive Health Measurement — chair: Richard Gershon | Theater |
| Computerized Adaptive Testing (CAT) for Patient-Reported Outcomes Measurement Information System (PROMIS): Evaluation of Stopping Rules for the Pediatric Item Banks (Jiwon Kim) | ||
| NIH Toolbox Calibration and Scaling for a multi-stage episodic memory test across the lifespan (Emily Ho) | ||
| CAT on Mobile Devices: Assessments of Verbal Ability for Remote Administration in Mobile Toolbox (Y. Catherine Han) | ||
| Individual Presentations: Applications of CAT — chair: Mariana Curi | Room 1110 | |
| An IRT-Based Adaptive Measure Of Conscientiousness In The Big Five Framework (Gabriela S. Lozzia) | ||
| Adaptive Test of Vocational Interests (TIPA): Development and psychometric properties of a computerized assessment measure (Gustavo Martins) | ||
| Using Item Response Theory to Develop a Parsimonious Measure of Psychological Need Satisfaction at Work (Alexsandro Luiz de Andrade) | ||
| Individual Presentations: Applications of CAT — chair: Roberta Palacios | Room 1150 | |
| Review and Requalification of a Workforce Competency Assessment Tool for Technical Education (Paloma de Lima Santos) | ||
| Design and Implementation of Practice-Oriented, Large-Scale Certification Exams: The Case of ANBIMA's Distribution Certifications (Andrea Burgos de Azevedo Mangabeira) | ||
| 12:00 - 14:00 | LUNCH | Free |
| 14:00 - 15:40 | Symposia: Advancing Adaptive Measurement: Dr. David Weiss's IRT and CAT Lab Innovations in CAT Methodology and Adaptive Measures of Change — chair: Robert Chapman | Auditorium 10th floor |
| Change Pattern Detection in Adaptive Measure of Individual Change (Raj Wahlquist) | ||
| Interactions Between Termination Criteria and Ability Estimators in Computerized Adaptive Testing (Xinyu Liu) | ||
| One of these things is not like the other: Setting Thresholds for Equivalence of Item Response Theory Parameter Estimation (Advancing Adaptive Measurement) (Robert Chapman) | ||
| Tracing the Roots and Branching into the Future: The Academic Legacy of Dr. David J. Weiss (Advancing Adaptive Measurement) (Robert Chapman) | ||
| Individual Presentations: MST, Item generation, item banking, form assembly, and others — chair: Heliton R. Tavares | Theater | |
| Automatic Pre-Testing of Mathematics Assessment Items: Predicting Item Response Theory (IRT) Difficulty Parameter with Machine Learning Methods (Patrick Canto de Carvalho) | ||
| A Stochastic Constrained Test Assembly Method Applied to Multistage Adaptive Testing (Alina von Davier) | ||
| Designing multi-stage CAT assessments of academic English language proficiency to give examinees the greatest opportunity to demonstrate that they meet performance level thresholds (Stephen Walker) | ||
| parATA: an R-package for Automated Test Assembly (Angela Verschoor) | ||
| Individual Presentations: Statistical methodology — chair: Ricardo Primi | Room 1150 | |
| Beyond Correct and Incorrect: A Nested Logit Computerized Adaptive Test Using Distractor Information to Assess Reading Proficiency (Thiago Costa) | ||
| Computerized Adaptive Testing Using Nonparametric Item Response Theory Models (Cecilia Marconi) | ||
| Addressing Misclassification Bias in Computerized Adaptive Testing with Latent Performance Standards (Bartosz Kondratek) | ||
| Testlet-Based Termination in Relaxed CAT: Reducing Assessment Length and Testlet Exposure (Luciana Cancado) | ||
| Computer Adaptive Testing with small item bank: Improved precision and reduced opportunities for cheating (Rense Lange) | ||
| 15:40 - 16:10 | COFFEE BREAK | Foyer Teatro |
| 16:10 - 17:10 |
Keynote 5: Ricardo Primi — chair: Mariana Curi Navigating the Psychometric AI Revolution: Innovations in Scoring and Validity |
Teatro |
| 17:10 - 17:30 | Ending Ceremony | |
Thursday-Friday, 1-2 October
- Organized tour in Rio de Janeiro*
*Additional Options
Pre-conference Workshops
Introduction to IRT and CAT
Nathan Thompson (ASC)
AI Applications in Adaptive Testing
Duanli Yan (Measurement)
Alina A. von Davier (Duolingo)
The Shadow-Test Approach to Adaptive Testing
Seung W. Choi (UT-Austin)
Richard J. Patz (University of California-Berkeley)
Wim J. van der Linden (University of Twente)
Introduction to IRT and CAT
Nathan Thompson (ASC)
New to CAT? This workshop is for you. We will start with an introduction to the CAT algorithm, including item selection, exposure constraints, scoring, and termination rules. We will discuss how this algorithm can be adapted to various approaches, including multistage testing, and how to utilize simulations to design and validate your CAT. Workshop assumes you are familiar with item response theory, and ready to apply it in adaptive testing.
AI Applications in Adaptive Testing
Duanli Yan (ETS)
Alina A. von Davier (Duoling)
In the era of artificial intelligence (AI), the field of educational testing faces significant challenges, particularly in test development and scoring in adaptive assessment, two key innovations include automated item generation (AIG) and automated scoring (AS). Only recently that generative AI has facilitated the development of complex test items on a large scale. We introduce “the item factory”, for managing large-scale test development including automation of item generation, quality review, quality assurance, and crowdsourcing techniques in adaptive testing. We present an overview of the latest natural language processing (NLP) techniques and large language models for AIG, alongside psychometric principles and practices for test development. We discuss the application of engineering principles in designing efficient item production processes (Luecht, 2008; Dede et al, 2018; von Davier, 2017). As AS becomes an integral part of the assessment landscape due to their advantages in reporting time, cost, objectivity, consistency, transparency, and feedback. We aim to demystify AS and provide a comprehensive understanding of its workings. We offer an overview of the design, development, evaluation, and quality control of automated scoring systems, along with practical advice and considerations for practitioners on the applications of these systems into formative and summative assessments (Yan, Rupp, & Foltz, 2020). We will share the most recent AI applications in educational learning and assessment based on our upcoming volume AI for Measurement in Educational Learning and Assessment (von Davier and Yan, 2026).
The Shadow-Test Approach to Adaptive Testing
Seung W. Choi (UT-Austin)
Richard J. Patz (University of California-Berkeley)
Wim J. van der Linden (University of Twente)
It is tempting to think of adaptive testing as a more complicated version of the problem of automated assembly of fixed test forms. But thanks to a simple twist of the latter, the problem is easier to solve, always produces maximum information about the ability of each of the test takers, meets any blueprint in force for the test, and easily generalizes to testing formats with varying degrees of adaptation such as item-level adaptive testing, linear-on-the-fly testing, standard multistage adaptive testing, and multistage testing with adaptive routing tests.
The course has three different parts. In the first part, we explain the ideas underlying the shadow-test approach, discuss a few practical aspects of its implementation, and show some of its generalizations to different test formats. The next two parts are to introduce two software packages available for the implementation of the shadow-test approach to adaptive testing, offering the participants hands-on experience with the R package TestDesign and Optimal CAT, a currently freely downloadable microservice available for easy integration with common test delivery systems.
Participants are expected to bring their own laptops to the course. Handouts will be sent to the participants prior to the conference to prepare for the course and review its content afterwards.
