# Video Lab: AI analysis of ethnographic and IHUT video

Video Lab is a [Cassi.ai](https://www.cassiai.com) platform that analyzes ethnographic videos and in-home product tests, such as in-home usage tests (IHUT), consumption diaries and unboxings, with computer vision and multimodal language models. It breaks each video down scene by scene, transcribes the speech with timestamps, identifies objects, packs and gestures, and crosses what the person says with what the person does. The output is a set of behavioral KPIs where every conclusion carries a timestamp and a frame as evidence.

## Fifty participants film three moments of use. Someone has to watch all of it.

Video is the closest a researcher gets to the real moment of use. It is also the material that takes longest to turn into a number.

- Topic: Ethnographic video and in-home product tests. In-home usage tests (IHUT), consumption diaries, unboxings. The participant films the real moment of use at home, and the recording shows what no questionnaire captures: the hand, the pack, the dose, the hesitation.
- How it is done today: Watching at 1.5x and filling a spreadsheet. A study with 50 participants recording 3 usage moments produces more than 25 hours of raw video. The team watches it sped up and logs each behavior by hand. Between the end of fieldwork and the executive report, more than 20 working days usually go by.
- The pain: Nobody has time to measure said against done. Because human logging is heavy, samples stay small, at 10 to 15 people. And the most valuable reading of the study, the gap between the behavior declared in the questionnaire and the behavior recorded on video, is left to impressions.
- Where Cassi.ai comes in: Every video read scene by scene, with evidence. Video Lab breaks each recording down scene by scene, transcribes the speech with timestamps, identifies objects, packs and gestures, and crosses speech with action. The result is behavioral KPIs where every conclusion points to a timestamp and a frame.

## What the recordings can answer once all of them have been watched.

- Said against done ("Do they use the product the way they say they do?"): The self-report and the physical evidence sit in the same table, behavior by behavior, with the size of the gap. The distance between questionnaire and video becomes a measured number.
- Timestamp and frame ("Where exactly in the video does that happen?"): Every conclusion links to the second of the recording and to the frame that supports it. Anyone in the meeting can open the evidence and check it.
- Objects, packs, gestures ("What did the participant actually pick up?"): Computer vision describes each segment: the pack on the counter, the dosing cap, the hand turning the clothes inside out. Brand and product identification goes through human review before it is treated as settled.
- Unknown, not visible ("What if the camera did not show it?"): A field the video cannot support is marked unknown, not visible or not applicable. The platform leaves the gap in the base and sends the uncertain item to a person.
- Dashboard and base ("Can I cut this by profile?"): The consolidated base feeds a multi-project dashboard with filters by profile, and exports as a tabulated file to CSV, Excel and SPSS for the cuts your team runs on its own.
- Clip library ("Can I show the clip to the board?"): The key moments become a library of anonymized clips, with faces protected by mosaic, ready for the executive presentation.

## Six steps from raw file to report, with a person reviewing what is uncertain.

1. Inventory the videos. Every file is listed and receives a fingerprint, a unique signature computed from its content, together with audio and technical metadata. No video is processed twice or left out.
2. Key frames and contact sheet. The platform extracts the key frames of each recording and lays them out on a contact sheet, the whole video at a glance. Every frame keeps its timecode.
3. Transcribe and observe. Speech is transcribed with timestamps. Each segment gets a visual observation: objects, packs, gestures and the order in which things happen. Speech and action share one timeline.
4. Extract the KPIs. Each video is converted into the same fixed structure of behavioral KPIs, so participants can be compared and tabulated. What is not visible is marked not visible.
5. Validate and review. Structure and confidence are checked for every output. Uncertain items go to a human review queue before they enter the base. Rule: 95% of KPI outputs must pass structure validation.
6. Consolidate and deliver. The reviewed outputs form one consolidated base, from which the insights, quotes, frames and the report are produced. No conclusion without timestamp and frame.

## Two numbers of the manual flow, one target and one rule.

- 25+ h: of raw video in a typical study with 50 participants recording 3 usage moments
- 20+ days: working days between the end of fieldwork and the executive report in the manual flow
- < 2 h: processing target for that same batch in Video Lab
- 95%: of KPI outputs must pass structure validation, the technical acceptance rule of the product

## What it never does

- Never invents what is not visible. The field is marked unknown, not visible or not applicable.
- Never treats a brand or product identification as perfect without human review.
- Never infers sensitive data that is not visible in the video.
- Never shows a conclusion without its timestamp and frame.
- Never publishes a clip without face protection.

## Who it is for

Insights directors, Research managers and analysts, Research firms, Agencies running qualitative video studies. For teams that already collect video and want the whole sample read with the same criteria, with the evidence one click away.

## Questions and answers

### What is Video Lab?

Video Lab is a Cassi.ai platform that analyzes ethnographic videos and in-home product tests with computer vision and multimodal language models. It breaks each video down scene by scene, transcribes speech with timestamps, identifies objects, packs and gestures, and crosses what the person says with what the person does.

### What kinds of video does it analyze?

Recordings made by participants at home or in context: in-home usage tests (IHUT), consumption diaries, unboxings and ethnographic videos in general. The typical case is a study in which each participant films several moments of use of a product.

### How does it compare what people say with what they do?

The transcription and the visual observation share one timeline. For each behavior, the platform records what was declared, in speech or in the questionnaire, and what the video physically shows, and reports the gap between the two with the timestamp and the frame of the evidence.

### What happens when something is not visible in the video?

The field is marked unknown, not visible or not applicable. Video Lab never fills a gap with a guess. Items with low confidence go to a human review queue before they enter the consolidated base.

### Does it identify brands and products on its own?

It detects packs and objects in the scene and suggests what they are. A brand or product identification is never treated as perfect without human review, so these items pass through a person before they reach the report.

### How long does the processing take?

A typical study with 50 participants recording 3 usage moments produces more than 25 hours of raw video, and the manual flow usually takes more than 20 working days up to the executive report. The processing target for that same batch in Video Lab is under 2 hours, followed by the human review of uncertain items.

### How is participant privacy handled?

Clips that leave the platform have faces protected by mosaic, and no clip is published without that protection. Video Lab also never infers sensitive data that is not visible in the video.

### What do I receive at the end?

An interactive executive presentation on the web with video frames, a multi-project dashboard with filters by profile, a tabulated base exportable to CSV, Excel and SPSS, a strategic report, and a library of anonymized clips.

## Bring a video study you know well.

Tell us about a study you are running, or one you have already read by hand. We show how Video Lab handles that kind of batch and walk you through a results deck, with the timestamp and the frame behind each conclusion. https://videolab.cassiai.com/#acesso
