Skip to content

Studies

Two studies are complete: an online pre-study that selected the interviewer avatars, and the VR interview study that evaluated the platform.

The instrument works, and it measures things a conventional survey cannot see. The behavioural effects it reveals are real but modest at n = 27, and are reported here with their confounds and their nulls intact.


Study 1: The avatar pre-study

Fielded 6–24 June 2025. 99 of 115 respondents completed it, rating 20 avatars on preference, trust, comfort and similarity. None took part in the interviews.

Four avatars went into the main study. They were not the four best-rated: they were chosen to span most of the preference range, so that avatar effects could be looked for at all.

preferencetrustcomfortsimilarity Avatar 3, sunglasses Sunglassespre-study no. 3
interview id 1
5 0 0 11 Avatar 8, headscarf Headscarfpre-study no. 8
interview id 2
42 63 68 0 Avatar 12, blue hair Blue hairpre-study no. 12
interview id 3
100 63 68 63 Avatar 11, striped jumper Striped jumperpre-study no. 11
interview id 4
53 42 42 79

Percentile among the twenty pre-study avatars, 0 to 100. The tick in each track marks the 50th percentile.

The four carried forward. Percentile among the twenty on each construct. The sunglasses avatar is last of twenty on both trust and comfort; the blue-haired avatar is the single most preferred of the field. Note that the high/low split holds on preference but not on the others: the lower-ranked headscarf avatar scores higher on trust and comfort than the higher-ranked striped jumper.

No causal claim about avatars

Interviewer avatar was not randomised across the 27 interviews, and the interviewee set is a different, unbalanced subset. The pre-study establishes how the avatars were perceived by an independent sample; it does not license a causal claim in the main study.


Study 2: The VR interview study

Twenty-seven interviews at two institutions. Interviewers and interviewees had never met. Each ran about 20 minutes of questionnaire inside a median 27.7-minute session, including two personality batteries used as the experimental contrast: neutral (QN) and discomfort-inducing (QD).

What the instrument captured

Records stored, by modality

motion capture speech

Effective gaze sampling rate, 27 interviewee streams

17 18 19 20 21 Hz median 18.5 Hz

One dot per stream, read back from timestamps rather than taken from a device datasheet.

Gaze yield against session duration

30k 40k 50k 60k 70k 20 40 60 80 session duration (min) P02: 30.1 min, 50,776 samples P03: 45.2 min, 71,200 samples P04: 34.5 min, 67,947 samples P05: 27.7 min, 51,930 samples P07: 18.1 min, 38,497 samples P08: 20.1 min, 41,369 samples P12: 43.4 min, 34,459 samples P14: 28.1 min, 56,418 samples P15: 19.5 min, 36,736 samples P16: 30.8 min, 36,340 samples P17: 36.0 min, 39,280 samples P18: 14.3 min, 28,226 samples P20: 22.4 min, 48,743 samples P21: 62.0 min, 37,839 samples P24: 25.1 min, 48,684 samples P25: 21.8 min, 44,238 samples P27: 28.1 min, 60,375 samples P28: 24.4 min, 50,363 samples P29: 39.7 min, 64,918 samples P30: 23.6 min, 45,082 samples P31: 18.9 min, 35,171 samples P32: 28.7 min, 58,215 samples P33: 32.3 min, 62,973 samples P34: 16.9 min, 33,312 samples P35: 25.8 min, 48,744 samples P36: 22.4 min, 41,005 samples P23: 203 min recorder left running

Gaze samples per session, both players combined.

Measured, not specified. Records per modality; the effective gaze sampling rate across 27 interviewee streams, read back from timestamps rather than taken from a device datasheet; and gaze yield against session duration.
27valid interviews, of 34 recorded
4,567,743records stored
15.7 htotal interview time
18.5 Hzmedian gaze rate
1,496,517eye, body and head records, each
0.966median capture duty cycle
57,971transcribed words
97.0 %gaze samples in the reference view

A duty cycle of 0.966 means capture is continuous for essentially the whole session. This is the claim everything else rests on, and the one we can make most firmly.

Acceptance

strongly disagreerather disagreepartlyrather agreestrongly agree
100 %50 %050 %100 %
Post-interview self-report. Discomfort and dizziness are at the floor (M = 1.48 and 1.39), willingness to take part again near the ceiling (M = 4.65), immersion moderate (M = 3.35). Crucially these are independent of the interviewer's avatar and of avatar–participant matching: the experience is stable across conditions, which is what makes the environment usable as a controlled setting.

Answer format

The cleanest result in the study, and the one most likely to matter elsewhere. The answer options were displayed on the desk panel, numbered, and never read aloud.

Bare index“three”36.3 %
Full scale label“rather good”32.6 %
Neitherparaphrase or other26.8 %
Bothindex and label4.4 %
0 %10 %20 %30 %40 %
How participants voiced a scale answer. Share of spoken answers to scale items, resolved against each item's actual response-option catalogue rather than a single hard-coded five-point vocabulary.

Two things follow. Since the indices were only ever visible, a participant could produce one only by reading the panel and folding it into their answer: the interface demonstrably changes response behaviour, and interviewers unanimously reported that displaying the options helped the interview flow.

And 63 % of spoken answers contain no verbatim scale label. A voice-driven questionnaire therefore cannot match on label text; it has to resolve a bare index against whatever scale is displayed, and fall back to a clarification turn for the 27 % that give neither form. That constraint generalises well beyond VR.

Index usage does not differ between neutral and sensitive items (W = 70.0, p = .758); it is a property of the interface, not of question sensitivity.

Gaze

Gaze is resolved by replaying captured eye tracking into the Unity scene and raycasting against the actual colliders, so "looked at the answer panel" is a measurement rather than an inference from a 2D projection.

Where people look barely moves. No area-of-interest dwell shift survives correction for multiple comparisons (n = 26, all Holm p = 1.0); the largest raw shift is toward the window (+0.026, p = .162), with the answer panel, self-view mirror and interviewer's face all essentially flat.

How people look does move.

Metric QN QD p Holm
Fixation time share (s/s) 0.628 0.583 .008 .040
Mean fixation duration (s) 0.338 0.289 .029 .117
Fixations per second 1.790 1.817 .940 1.0
Gaze dispersion (deg) 13.14 14.61 .745 1.0

Fixation time share is the one metric that survives Holm correction across the fixation family (dz = −0.64): under sensitive questions participants spend a smaller fraction of their time in fixation, shifting toward slightly longer, less frequent fixations. Spatial entropy is also higher under sensitive items (2.98 → 3.17, p = .018, n = 20): participants scan more of the scene rather than staring more widely.

Response latency

Sensitive items are answered 0.42 s more slowly at the median (2.73 s vs 2.31 s), but this is not a reliable effect and the design cannot make it one (Mann-Whitney p = .210).

The confound, stated plainly

The batteries are not interleaved: every neutral item precedes every sensitive one, so item identity explains 99.3 % of the variance in interview position. Question type, position and item are very nearly the same variable. Under participant-clustered inference neither term is reliable: sensitivity ×0.89 (p = .613), position ×1.48 (p = .143), with the model's sensitivity estimate pointing in the opposite direction from the raw comparison. We therefore claim no latency effect of either kind. Interleaving the batteries is the fix, and belongs in the next study.

Speech shows the same shape: responses to sensitive items run longer (1.51 s → 2.02 s, p = .025), but the effect does not survive correction.

Behaviour against self-report

The finding that most directly justifies building the instrument.

Asked directly afterwards, only 10 of 27 respondents (37 %) said any question had made them uncomfortable. Those 10 differ sharply from the other 17:

Measure Claimed discomfort Did not d Holm p
Δ response duration +1.00 s 0.00 s 1.76 .0014
Δ interviewer-face dwell +0.005 0.000 0.73 .037
Δ median latency, latency ratio, Δ mirror dwell n/a n/a n/a n.s.

Both surviving effects are substantial. But the crucial observation is the other half: the same behavioural shifts appear in participants who reported no discomfort at all. Self-report and behaviour converge only partially, so if those shifts index discomfort, self-report alone substantially underestimates the impact of sensitive questions, plausibly through social-desirability reluctance. That is precisely the gap a multimodal interview environment exists to close.

A separate analysis, not to be confused with this one

Correlating behavioural shifts against the five post-interview Likert items gives 25 rank correlations, none of which survive Holm correction. At n = 27 the intervals are wide enough to contain moderate effects, so that is an absence of demonstrated convergence, not demonstrated independence: a different question from the group comparison above.


Reproducibility

Every figure and number here comes from an analysis package that regenerates end to end in about a minute (python run_all.py). Only the cache-building step touches the network; everything downstream is offline and deterministic, and a clean rebuild reproduces every reported number exactly. The package's summary document is written by reading results back out of the generated tables, so the prose cannot drift from the data.

Limitations

  • n = 27. Intervals are wide; directional results should be read as directional.
  • Fixed battery order makes question type and time-on-task inseparable here.
  • Interviewer avatar was not randomised, so no causal avatar claim is made.
  • Transcription is imperfect, which affects the "neither" answer-format category and under-detects filled pauses.
  • Interviewer gaze is captured but not analysed. It projects into the interviewee's reference view for a median of about 1 % of samples, and needs its own reference geometry.

Full detail is in the InterView paper.

Partners & funding