Human Interaction Testing Platform | Bingham Intelligence
Product

From live interaction to usable data.

Define the scene. Capture the interaction. Review what happened. Turn the session into structured labels, behavioral findings and data your team can use.

One workflow, from brief to score.

Performers interact with your live product. Scene instructions, cue logs and recordings connect each test to what actually happened.

SCENE 07 OF 20 · PAUSE HANDLING
Ask the agent for a restaurant tip. Pause mid-sentence to think.
Variation B: look up and away for two seconds when cued.
Use your own words. Scenes with set lines show them here.
Start session
01
BriefThe performer learns the situation and goal, plus exact lines when a scene needs them.
● RECORDING00:03.05
Your AI agent, live
Cue: look away now
Also played in your earpiece
02
PerformThey meet your live agent. The app cues the key moment at an exact time.
● CAPTURING
Performer video1080p, 30 fps
Performer audio48 kHz
Agent outputas streamed
Cue logmilliseconds
Agent eventsfrom API
Every stream shares one clock.
03
CaptureVideo, audio, cues and your agent's API events land on one clock.
SESSION S0007 · LABELED
Pause handling
Cue fired3.05 s
Agent API event3.42 s
FAIL
Review the recording against the agreed criteria.
04
ScoreLabels connect observed behavior to agreed criteria. Human review checks judgments before retesting your next release.
Interaction workflow

See how an interaction becomes evidence.

Four interaction scenarios show how timing and nonverbal behavior can be tested. Studies also cover emotion, role consistency, repair and collaboration.

Pause handlingFAIL

Ask the agent for a restaurant tip. Pause mid-sentence to think.

LIVE CUELook up and away, hold the pause for two secondsMEASUREDoes the agent wait while the person is visibly thinking?PASS MEANSThe agent stays silent until the person finishes the sentence.RESPONSEAgent API event: 0.330 s after the observed pause began
Nonverbal interruptionFAIL

Let the agent explain its return policy, then stop it without words.

LIVE CUERaise a flat hand toward the cameraMEASUREDoes the agent stop speaking within 1.5 seconds?PASS MEANSThe agent yields within 1.5 seconds of the gesture.RESPONSEAgent stopped 2.4 s after the raised hand
BackchannelingPASS

Listen to the agent give directions and show you are following.

LIVE CUENod twice and say "mm-hm"MEASUREDoes the agent keep going instead of stopping to reply?PASS MEANSThe agent keeps its turn through the nod.RESPONSEAgent kept its turn through the backchannel
Gaze aversionPASS

Describe a problem with your order while searching for the right word.

LIVE CUEGlance down and to the side mid-sentenceMEASUREDoes the agent read the glance as thinking, not finished?PASS MEANSThe agent waits until the sentence ends.RESPONSEAgent waited 1.1 s past the glance
Session format and scoring.

Capture the event. Verify the behavior.

Cue logs record the intended action. Recording review confirms what the performer actually did. Observed behavior and agent events are aligned before labels are assigned. Training collections and held-out evaluation sets remain separate.

FILEWHAT IT HOLDSRATE sessions.csvOne row per session: scene, variation, agent build, consent scopeper session performer.mp4Front camera video of the performer1080p, 30 fps performer.wavPerformer audio on its own channel48 kHz agent.mp4, agent.wavThe agent's output exactly as receivedas streamed cues.jsonlEvery cue the app firedmillisecond timestamps agent_events.jsonlAgent speaking start and stop, from its APImillisecond timestamps transcript.jsonlWords with start and end times, both sidesper word vad.jsonlVoice activity, both sides100 Hz keypoints.npzFace, hand and body landmarks30 Hz sync.jsonClock alignment detailsper session labels.jsonEvent window, dynamic, result, rater scoresper event
SESSION S0007_P012 · PAUSE HANDLING · VARIATION B
0.000appsession_startScene S07, variation B
1.200performerspeech_start"So I was wondering if you know a place"
3.050appcue_firedLook away, thinking pause, 2.0 s
3.090performerspeech_stopPause begins
3.420agentstarted_speakingFrom the agent's API event
5.100performerspeech_start"that's open late?"
5.250agentstopped_speakingInterrupted: true
6.900agentstarted_speakingAnswers the question
Result: FAIL. The agent API event occurred 0.330 s after the observed pause began.
{
  "session_id": "S0007_P012",
  "dynamic": "pause_handling",
  "event_window": { "t_start": 3.090, "t_end": 5.100 },
  "cue": { "type": "look_away_pause", "fired_at": 3.050 },
  "agent_api_event_s": 3.420,
  "delay_from_observed_pause_s": 0.330,
  "delay_from_cue_s": 0.370,
  "result": "fail",
  "consent_scope": [ "evaluation", "training" ]
}
Seamless Interaction-style foldersVideoFDB-style event windowsRobot-specific exports scoped with partnersHugging Face datasets
Session S0007. API events are checked against recordings before reporting audible response timing. Timestamp precision does not establish measurement accuracy.

Under the hood.

Our workflow connects scene design, capture and human review. The result is event-level evidence your team can inspect and a protocol you can repeat after an update.

Scene formatA script a computer can read: situation, cues, timing and what counts as a pass. Every performer and system follows the same notes.
CaptureA flight recorder for the session: both sides, every cue and the AI's own events on one shared clock.
Assisted labelingCombine event timestamps with recording review. Apply agreed criteria and flag judgments that need another reviewer.
ConvertersStructured exports matched to the customer’s agreed schema and use case. Robot-specific formats require integration and validation.
Benchmark harnessA standardized test center: the same exam for every system, with grades and a full report.

Consent built in.

Every performer gives separate consent for testing, publication and training. Every delivery includes a rights record for each session.

Ready to test your AI with real people?

Start with an Interaction Study on your live product.

Book a study