Skip to content
cassi.aiVideo Lab
Request access

00:00:00 · Automated analysis of ethnographic video

Every video in the study, read in depth by AI. Evidence behind every conclusion.

A whole batch of in-home test recordings comes in at once and comes out read scene by scene: speech transcribed, objects, packs and gestures identified, and what each person says placed next to what she does. The participant says one cap and pours a cap and a half. Every conclusion carries a timestamp and a frame.

Illustrative data. The scenes play on their own.

00:01:00 · The work

Fifty participants film three moments of use. Someone has to watch all of it.

Video is the closest a researcher gets to the real moment of use. It is also the material that takes longest to turn into a number.

Topic

Ethnographic video and in-home product tests

In-home usage tests (IHUT), consumption diaries, unboxings. The participant films the real moment of use at home, and the recording shows what no questionnaire captures: the hand, the pack, the dose, the hesitation.

How it is done today

Watching at 1.5x and filling a spreadsheet

A study with 50 participants recording 3 usage moments produces more than 25 hours of raw video. The team watches it sped up and logs each behavior by hand. Between the end of fieldwork and the executive report, more than 20 working days usually go by.

The pain

Nobody has time to measure said against done

Because human logging is heavy, samples stay small, at 10 to 15 people. And the most valuable reading of the study, the gap between the behavior declared in the questionnaire and the behavior recorded on video, is left to impressions.

Where Cassi.ai comes in

Every video read scene by scene, with evidence

Video Lab breaks each recording down scene by scene, transcribes the speech with timestamps, identifies objects, packs and gestures, and crosses speech with action. The result is behavioral KPIs where every conclusion points to a timestamp and a frame.

The questionnaire records what people declare. The video records what they do. The finding lives in the distance between the two.

00:02:00 · What it answers

What the recordings can answer once all of them have been watched.

Each part of the platform exists because a researcher asks this of a video study before signing the report.

"Do they use the product the way they say they do?"

Said against donedeclared and observed

The self-report and the physical evidence sit in the same table, behavior by behavior, with the size of the gap. The distance between questionnaire and video becomes a measured number.

"Where exactly in the video does that happen?"

Timestamp and frameper conclusion

Every conclusion links to the second of the recording and to the frame that supports it. Anyone in the meeting can open the evidence and check it.

"What did the participant actually pick up?"

Objects, packs, gesturesper segment

Computer vision describes each segment: the pack on the counter, the dosing cap, the hand turning the clothes inside out. Brand and product identification goes through human review before it is treated as settled.

"What if the camera did not show it?"

Unknown, not visibleskepticism rule

A field the video cannot support is marked unknown, not visible or not applicable. The platform leaves the gap in the base and sends the uncertain item to a person.

"Can I cut this by profile?"

Dashboard and baseCSV, Excel, SPSS

The consolidated base feeds a multi-project dashboard with filters by profile, and exports as a tabulated file to CSV, Excel and SPSS for the cuts your team runs on its own.

"Can I show the clip to the board?"

Clip libraryfaces under mosaic

The key moments become a library of anonymized clips, with faces protected by mosaic, ready for the executive presentation.

00:03:00 · Method

Six steps from raw file to report, with a person reviewing what is uncertain.

The same sequence runs for every video, so participant 3 and participant 47 are read with the same criteria. For a typical batch of 50 participants, the processing target is under 2 hours.

  1. 01

    Inventory the videos

    Every file is listed and receives a fingerprint, a unique signature computed from its content, together with audio and technical metadata.

    No video is processed twice or left out.

  2. 02

    Key frames and contact sheet

    The platform extracts the key frames of each recording and lays them out on a contact sheet, the whole video at a glance.

    Every frame keeps its timecode.

  3. 03

    Transcribe and observe

    Speech is transcribed with timestamps. Each segment gets a visual observation: objects, packs, gestures and the order in which things happen.

    Speech and action share one timeline.

  4. 04

    Extract the KPIs

    Each video is converted into the same fixed structure of behavioral KPIs, so participants can be compared and tabulated.

    What is not visible is marked not visible.

  5. 05

    Validate and review

    Structure and confidence are checked for every output. Uncertain items go to a human review queue before they enter the base.

    Rule: 95% of KPI outputs must pass structure validation.

  6. 06

    Consolidate and deliver

    The reviewed outputs form one consolidated base, from which the insights, quotes, frames and the report are produced.

    No conclusion without timestamp and frame.

00:04:00 · The numbers

Two numbers of the manual flow, one target and one rule.

The first two describe how the work is done today. The last two state how Video Lab is built to work. None of them is a promise about your study.

25+ hof raw video in a typical study with 50 participants recording 3 usage moments
20+ daysworking days between the end of fieldwork and the executive report in the manual flow
< 2 hprocessing target for that same batch in Video Lab
95%of KPI outputs must pass structure validation, the technical acceptance rule of the product

What it never does

  • Never invents what is not visible. The field is marked unknown, not visible or not applicable.
  • Never treats a brand or product identification as perfect without human review.
  • Never infers sensitive data that is not visible in the video.
  • Never shows a conclusion without its timestamp and frame.
  • Never publishes a clip without face protection.

Who it is for

  • Insights directors
  • Research managers and analysts
  • Research firms
  • Agencies running qualitative video studies

For teams that already collect video and want the whole sample read with the same criteria, with the evidence one click away.

00:05:00 · Questions

Questions and answers

What is Video Lab?

Video Lab is a Cassi.ai platform that analyzes ethnographic videos and in-home product tests with computer vision and multimodal language models. It breaks each video down scene by scene, transcribes speech with timestamps, identifies objects, packs and gestures, and crosses what the person says with what the person does.

What kinds of video does it analyze?

Recordings made by participants at home or in context: in-home usage tests (IHUT), consumption diaries, unboxings and ethnographic videos in general. The typical case is a study in which each participant films several moments of use of a product.

How does it compare what people say with what they do?

The transcription and the visual observation share one timeline. For each behavior, the platform records what was declared, in speech or in the questionnaire, and what the video physically shows, and reports the gap between the two with the timestamp and the frame of the evidence.

What happens when something is not visible in the video?

The field is marked unknown, not visible or not applicable. Video Lab never fills a gap with a guess. Items with low confidence go to a human review queue before they enter the consolidated base.

Does it identify brands and products on its own?

It detects packs and objects in the scene and suggests what they are. A brand or product identification is never treated as perfect without human review, so these items pass through a person before they reach the report.

How long does the processing take?

A typical study with 50 participants recording 3 usage moments produces more than 25 hours of raw video, and the manual flow usually takes more than 20 working days up to the executive report. The processing target for that same batch in Video Lab is under 2 hours, followed by the human review of uncertain items.

How is participant privacy handled?

Clips that leave the platform have faces protected by mosaic, and no clip is published without that protection. Video Lab also never infers sensitive data that is not visible in the video.

What do I receive at the end?

An interactive executive presentation on the web with video frames, a multi-project dashboard with filters by profile, a tabulated base exportable to CSV, Excel and SPSS, a strategic report, and a library of anonymized clips.

Bring a video study you know well.

Tell us about a study you are running, or one you have already read by hand. We show how Video Lab handles that kind of batch and walk you through a results deck, with the timestamp and the frame behind each conclusion.

It goes straight to the Cassi.ai team and is not added to any mailing list.