TARDIS: Bridging the Social Gap with AI Recruiters and Open Learner Models
Adolescents’ Self-regulation During Job Interviews Through an AI Coaching Environment
The paper introduces TARDIS, an intelligent AI coaching environment designed to improve job interview skills for at-risk adolescents (NEETs). By combining an embodied AI recruiter simulator with an Open Learner Model (OLM) called NOVA, the system provides situated rehearsals and data-driven reflection to enhance both verbal and non-verbal social competencies.
TL;DR
Job interviews are high-stakes social interactions that require intense self-regulation. This paper presents TARDIS, an AI-driven coaching environment that uses virtual recruiters and data visualization (Open Learner Models) to help adolescents master these skills. The research proves that practicing with AI agents leads to superior improvements in eye contact, tone of voice, and answer quality compared to traditional web-based training.
Background & Positioning
Developing social skills isn't about being told what to do; it's about doing it, reflecting, and adjusting. While mock interviews with humans are great, they are hard to scale and often lack objective data for reflection. TARDIS enters the scene as a situated learning tool that sits between theory and real-life practice, providing a repeatable bridge for "at-risk" youth (NEETs) to build confidence and competence.
The Problem: The Feedback Gap in Social Training
Traditional social skills interventions suffer from two main issues:
- Vicarious Learning (The "Watch and Learn" Trap): Reading social stories or watching videos provides theory but no muscle memory.
- Role-Playing Limitations: Human-to-human mock interviews are fleeting. Without a recording or objective data, the feedback is subjective and often fails to trigger the "Aha!" moment needed for self-regulation.
For adolescents dealing with anxiety and the threat of unemployment, the cost of failure in a real interview is too high. They need a "flight simulator" for social interaction.
Methodology: The TARDIS Architecture
The TARDIS environment is a sophisticated blend of affective computing and pedagogical scaffolding. It consists of two main pillars:
1. The Interaction Simulator
Using the FAtiMA planning architecture, the system generates AI recruiters (AIRs) with specific emotional intelligence levels. These aren't just chatbots; they process:
- Verbal Input: Dialogue via headsets.
- Non-Verbal Input: Gestures and posture via Microsoft Kinect.
This creates a "Face-Threatening" or "Understanding" environment that forces learners to regulate their physiological and emotional responses in real-time.
2. Open Learner Modelling (OLM) with NOVA
This is the "brain" of the feedback loop. Post-interview, the NOVA tool synchronizes video of the learner with raw data (pitch, amplitude, eye gaze).
The image shows the dual-interface: The AIR simulator (left) and the NOVA reflection tool (right).
Experimental Results: Moving the Needle
The study compared TARDIS against a standard UK Job Centre web-based program. While both groups felt more confident (self-efficacy), the objective AI-coached group outperformed the control group in crucial areas:
- Non-Verbal Mastery: Significant gains in Eye Contact (F=14.07) and Tone of Voice (F=13.88). The ability to see their own monotony or lack of gaze in the OLM data allowed learners to self-correct more effectively.
- High-Pressure Questions: Learners improved significantly more on "Why should we hire you?"—the ultimate test of self-presentation.
| Metric | Intervention Group (IG) vs Control (CG) | Significance (p) |
|---|---|---|
| Eye Contact | Much Higher Improvement | p < .01 |
| Tone of Voice | Much Higher Improvement | p < .01 |
| Facial Expression | Higher Improvement | p < .05 |
Deep Insight & Conclusion
The true power of TARDIS isn't just the "AI recruiter"—it's the Open Learner Model. By making "invisible" social behaviors (like a shaky voice or shifting eyes) visible through data, the system empowers the learner to engage in metacognitive reflection.
Limitations & Future Work
The study noted that "Tone of Voice" and "Eye Gaze" are culturally dependent and notoriously difficult for even human annotators to agree on (moderate Kappa scores). Future iterations aim to automate these annotations even further to provide real-time, objective feedback without needing a human practitioner to mediate every session.
Takeaway: AI in education isn't about replacing the teacher or the experience; it's about providing the situated data that makes human reflection 10x more effective.
