MMy First Million
← All frameworks
InnovationEmmett Shear

The Video System Of Record Inversion

Find consumer incumbents whose moat is a hand-typed database, and invert it so raw video becomes the record and AI extracts the fields.

Difficulty
Advanced
Time to result
~months to results
Steps
6
Confidence
80%

Shear's model is that a huge number of consumer businesses are effectively a database with a system of record - canonical rows about the world, with the messy real world digested into a row by users at write time. Yelp is his teaching example: rows are restaurants, fields are location and hours, and photos are attached facts about the row rather than the basis of it. The inversion is to make raw video the system of record and have AI extract a cached version of the metadata. The payoff is retroactive schema changes: if you later decide noise level matters, you do not run a new data collection process, you re-watch the existing videos - or let the user ask in real time and sort search results by a field that was never in the database. Shear thinks the resulting experience starts with the camera open, which is not Yelp today. He is careful about where it does and does not work: he calls Yelp a bad example because its photos and reviews are already about 90% as good as video would be, and notes some incumbents will just bolt on video and win. But where the text version is a poor compression of reality, you can both disrupt the incumbent and 10x the segment. He offers it as a general-purpose consumer move applicable to any product where you fill out forms, comparable to build-it-for-mobile.

Origin

Shear developed it while assessing what is genuinely new for consumer startups in the AI wave, using Yelp as the familiar case for how user-generated content apps really work.

Core principles

  • 01Most consumer apps are a database of canonical rows about the world
  • 02Users currently do the compression work at write time
  • 03AI moves compression from write time to read time
  • 04Retroactive schema changes are the real unlock
  • 05The bigger the loss in the text compression, the bigger the opportunity

How to run it

  1. 1

    Model the target as rows and fields

    Describe the incumbent as a system of record: what is the row, what are the canonical facts on it. For Yelp the rows are restaurants and local businesses and the facts are things like location and hours.

    Pro tip Any time you type into a text box, you are participating in one of these systems of record.

  2. 2

    Locate where the compression happens

    Establish that users go out into the messy world and turn it into a row at write time, with photos and video attached as facts about the row rather than forming the record itself.

    Watch out If the real asset is licences or supply rather than user data entry, this model does not apply.

  3. 3

    Score how lossy the text version is

    Judge how much of the underlying reality survives the compression to text. Shear rates Yelp's photos plus reviews at roughly 90% as good as a video record, which is precisely why he calls it a weak target.

    Pro tip Hunt for categories where it is not even obvious what should go in the processed text version.

  4. 4

    Invert to raw capture

    Store the raw video - the meal, the person talking about it, whether they had a good time - and have AI extract a cached metadata version from it. Design the experience accordingly, which Shear thinks probably starts with the camera open.

    Watch out A camera-first version of a familiar product will feel wrong at first; that is the point of the disruption.

  5. 5

    Prove the retroactive field

    Pick a field nobody collected - Shear's example is noise level - and extract it from the existing corpus on demand, letting users sort search results by it even though it was never in the database.

    Pro tip This is the capability the incumbent cannot match without re-running data collection.

  6. 6

    Check the disruption actually lands

    Test whether the incumbent can simply add videos and keep winning. Where the video record is genuinely better, Shear argues you both level the playing field and can 10x the size of the segment.

    Watch out It will not disrupt everything - in many categories the incumbent bolting on video is enough.

In the wild

A camera-first Yelp

Instead of typing a review, you upload raw video of the meal and of you and a friend talking about it. AI watches it and caches the metadata. Later someone asks what the noise level is at a restaurant, and the AI re-watches the videos in parallel and sorts fifteen results by a field that was never collected.

Yelp's meticulously groomed structured data stops being a moat, the playing field levels between the startup and the incumbent, and the entry point of the product becomes the camera.

ChatGPT against Google's ranked index

Shear's live example of the same inversion already playing out: Google's value was ranking web pages by terms so you could find a link to an answer. ChatGPT changed the unit from finding to answering, and then to executing a command that creates something rather than retrieving it.

The incumbent's structured understanding of the web stopped being the deciding asset, demonstrating the pattern working at full scale.

Common mistakes

Picking a target whose text record is already good enough

Shear calls Yelp a bad example on reflection, rating its photos and reviews at about 90% as good as a video system of record. If the compression loses little, the inversion buys you almost nothing.

Assuming the incumbent cannot respond

In many categories the incumbent can add videos, the result is not really better, and the incumbent simply wins. The inversion does not disrupt everything - the category has to be chosen for it.

Applying the model where the moat is not data entry

Shear rejects Spotify as an example: its value is the licences and the music itself, and if the human data entry parts of the database were deleted tomorrow it would barely hurt. Misidentifying the moat targets the wrong asset entirely.

Is it for you?

Best for

Consumer founders hunting for a category where AI genuinely levels the field against an entrenched incumbent

Not ideal for

B2B SaaS, and categories where the incumbent's structured data is already near-complete or the moat is licensing rather than data entry

From the transcript

a huge number of businesses can be conceived of as effectively being a database with a system of record

Emmett Shear · 26:00

you can take that and you can reapply it to any product where you fill out forms

Emmett Shear · 29:30

And so suddenly the playing field is leveled between the startup and Yelp, and that's a that's a huge opportunity for disruption.

Emmett Shear · 29:30

From the episode

Emmett Shear: Life After Twitch, Jeff Bezos Lessons & AI Doomsday Odds

Emmett Shear