Skip to content

Deliverables

Everything FACES has built is open source and public. Four repositories, each with its own documentation site; this page explains how they fit together.

Component What it is Repository Docs
InterView The interview platform. A Docker Compose stack that runs the whole system. texttechnologylab/InterView Docs · here
Va.Si.Li-Lab The Unity VR framework InterView's headset client is built on. texttechnologylab/Va.Si.Li-Lab Docs · here
Va.Si.Li-Lab-backend Scene, level and role definitions, the logging API, the Ubiq room server, and optional chatbot and speech services. texttechnologylab/Va.Si.Li-Lab-backend Docs · here
Janus-Gateway A pinned, reproducible container build of the Janus WebRTC server. texttechnologylab/Janus-Gateway Docs · here

How the pieces fit together

flowchart LR
  subgraph clients [Clients]
    direction TB
    VR["InterView-VR<br/>Unity, Meta Quest"]
    WEB["InterView-W<br/>browser"]
    LLM["LLM client"]
  end

  subgraph backend [InterView-B, the backend]
    direction TB
    JANUS["Janus SFU<br/>+ coturn TURN relay"]
    VSL["Va.Si.Li-Lab server<br/>Ubiq rooms and state"]
    QUEST["Questionnaire<br/>LimeSurvey"]
  end

  DB[("Logging API<br/>database")]

  subgraph proc [InterView-P, processing]
    direction TB
    ASR["ASR, Whisper"]
    GAZE["Gaze AOI replay"]
  end

  REL["Relational dataset"]
  DUUI["DUUI<br/>distributed processing"]

  VR --> JANUS
  WEB --> JANUS
  LLM -.-> JANUS
  VR --> VSL
  WEB --> VSL
  VSL --> QUEST
  JANUS --> DB
  VSL --> DB
  QUEST --> DB
  DB --> ASR
  DB --> GAZE
  ASR --> REL
  GAZE --> REL
  REL --> DUUI

  classDef planned stroke-dasharray: 5 4,stroke:#E5007D,color:#E5007D;
  class LLM planned;

Clients of any type join a room. Janus negotiates the audio and video streams between them and carries a control data channel; the Va.Si.Li-Lab server synchronises the environment, the avatars and the state of the questionnaire. Everything either of them produces is written through the logging API. InterView-P then collects each completed interview, transcribes the audio, resolves gaze against the geometry of the scene, imports the questionnaire responses, and links them all into one relational dataset, which can be handed on to DUUI for distributed multimodal processing.

The LLM client is drawn dashed because it is designed and supported by the backend but was not part of the evaluation study. See requirement (G) on the Project page.

The data model

The point of the architecture is what comes out of it. Interview data is not stored as isolated question–answer pairs with metadata attached; questionnaire items, multimodal recordings and contextual signals are nodes, and their temporal and behavioural relations are links.

The InterView data model Thirteen stored tables in two databases, 4,567,743 records in total. The timeseries database holds Experiment, Player, the per-frame Eye, Body and Head streams, and Audio, AudioChunk and the Words transcribed from them. The base database holds the questionnaire: QuestionWindow with its InterviewItem catalogue, and SurveyResponse and SurveyAnswer with the SurveyItem catalogue. QuestionWindow and SurveyResponse link back to Experiment across the two databases. TIMESERIES DATABASE captured on the session clock 4,562,480 records BASE DATABASE questionnaire and self-report 5,263 records cross-database links SESSION Experiment 34 SESSION Player 68 PER-FRAME Eye 1,496,517 PER-FRAME Body 1,496,517 PER-FRAME Head 1,496,517 MEDIA Audio 103 MEDIA AudioChunk 14,753 LANGUAGE Word 57,971 SELF-REPORT SurveyResponse 27 SELF-REPORT SurveyAnswer 3,537 INSTRUMENT SurveyItem 131 INSTRUMENT QuestionWindow 1,512 INSTRUMENT InterviewItem 56 KEY has many transcribed into records stored, log scale, 1 to 10⁷ Eye, body and head: one record per frame, per player
The data model, with the records actually stored. 4,567,743 records in thirteen tables across two databases, from the 27 interviews of the evaluation study. Everything captured on the session clock lives in the timeseries database: eye, body and head at one record per frame (1,496,517 each), and 57,971 words transcribed from 14,753 audio chunks. The questionnaire lives in the base database, and two links cross over: 1,512 question windows place each item on the interview timeline, and 27 survey responses (3,537 answers) tie self-report to the same session. Every count is read out of the live database, not estimated.

Because the modalities share a clock and a set of links, questions can be asked across them that no single stream could answer on its own: what a respondent was looking at while answering a particular item, how long they took, whether their fixation pattern shifted, and whether any of that lines up with what they later said about the experience. Studies shows what that makes visible.

A session timeline showing the separate data streams aligned on one clock

One clock. The separate streams of a single session, time aligned. Cross-modal synchronisation is what makes the relational structure usable rather than merely declared.

Reuse

The stack is designed to be run by other people. A minimal deployment is the database API, the Ubiq room server and a database instance; the full interview stack adds Janus, a TURN relay and the web client, and comes up with docker compose up -d once configured.

Read the setup guide before deploying

The stack requires an external database and a TLS-terminating reverse proxy, and will not work without them. See the InterView setup guide.

If you use any of it, please cite it.

Partners & funding