Recipe Book is a video data platform where you can semantically search 25M+ clips, iterate on the results, and buy the exact dataset you need to train your model in one sitting. Search -> rate a few clips -> we train a probe on your votes to re-rank the whole catalog to your taste. No contact forms, no sales calls, $3/hour.
What's up PH, I'm Braiden, co-founder of Shofo (YC W26).
Every team training a model faces the same data decision: build or buy. Building means spending engineering talent on collection (reverse engineer APIs, deal w proxies, fight scraping defenses, filtering, maintenance) all to collect data sitting in the open. Buying isn't much better. Prices run from $15 to $480 per hour, and your ideal dataset still sits behind a wall of sales calls, samples, and weeks (sometimes months) of collection and iteration. That time can be the difference between SOTA and irrelevant.
RecipeBook gives you the ability to search, filter, and build training datasets from millions of publicly available videos in the same day. You search in plain language, narrow with metadata filters (fps, resolution, duration, aspect ratio), and upvote and downvote clips based on needs. Once you vote, we train a classifier on your picks to re-rank the whole catalog to your taste, so you get the exact dataset you need. (Just a few votes can meaningfully change results)
Once you're happy with results, you type in how many hours you want. Ask for 1,000 and you get a review sheet: the score distribution, a playable grid of the worst clips in the order so you can see the floor, and a trim slider where the count and price all live. Check out at $3/hour and a manifest CSV with download links hits your inbox within a couple minutes (or instantly depending on size of purchase).
Biggest downfall is that metadata isn't super strong right now. No captions or engagement on the corpus (yet), so search is purely visual. (If you think it's something else please tell me!!)
Our team has spent years in data collection and curation. We've indexed billions of videos across the internet and built systems to serve that data on request across all modalities. Along the way we learned the data you need to train your model is already public, you just haven't found it yet.
If you're training a video or multimodal model, give it a try and tell me what you think.
If you know a team sourcing large-scale video data (video gen, VLMs, world models, seed to Series A), intro us and I'll buy you dinner.
Report
@braiden_dishman1 Congrats on shipping this. The volume of data processed is impressive
the search-then-buy flow is clever, but the thing i'd want answered before using this commercially is rights, not workflow. "publicly available" and "cleared for training a model you're going to sell" are two very different bars, and that gap is exactly what's driving a bunch of the current lawsuits against scrapers. is there any licensing/consent layer behind the 25M clips, or is it on the buyer to figure out whether a given clip is actually safe to train on?
@omri_ben_shoham1 Hey Omri, when we collect data we do so in a way that does not require logins, agreeing to TOS, creating fake accounts, etc. Our bar for collection methods is quite high and our approach is supported by public data case law. There are terms on RecipeBook's website that provide a bit more detail as well as a takedown form if that's helpful!
Report
@omri_ben_shoham1@braiden_dishman1 appreciate the direct answer, but I think it sidesteps the actual question. "we didn't need a login to get it" is a defense of how the data was collected, not a defense of what the buyer is allowed to do with it afterward. those are genuinely separate legal questions, and a takedown form only helps the original creator, it doesn't tell me, the buyer, whether the dataset I just bought is safe to train a commercial model on without getting pulled into someone else's lawsuit later. is there anything in the terms that actually indemnifies the buyer, or does the risk sit entirely on whoever purchases the hours
@omri_ben_shoham1@galdayan Self serve purchases don't come with indemnity coverage. Downstream risk sits with the buyer. (Similar structure most data providers use) On larger deals we can negotiate custom contracts with different reps and discuss indemnity. If you're interested happy to have that conversation.
Report
@omri_ben_shoham1@braiden_dishman1 appreciate the straight answer, that's more honest than most vendors would give. but it points at an odd shape: the buyers who can afford to negotiate a custom contract with indemnity are the ones who could probably absorb the legal risk anyway, and the self-serve buyer, the solo dev or small team paying $15-480/hour, is the one least equipped to deal with a downstream claim and also the one stuck with zero coverage by default. is a self-serve indemnified tier ever on the roadmap, or is the plan to keep that protection sales-gated indefinitely
Report
Searched for a tricky cooking motion I'd been hunting for and the re-rank after a few votes actually nailed it. The no-sales-call, pay-by-hour setup is refreshing.
RecipeBook by Shofo
@braiden_dishman1 Congrats on shipping this. The volume of data processed is impressive
RecipeBook by Shofo
@dmitrii_volosatov Thanks Dimitrii! More data is on the way
the search-then-buy flow is clever, but the thing i'd want answered before using this commercially is rights, not workflow. "publicly available" and "cleared for training a model you're going to sell" are two very different bars, and that gap is exactly what's driving a bunch of the current lawsuits against scrapers. is there any licensing/consent layer behind the 25M clips, or is it on the buyer to figure out whether a given clip is actually safe to train on?
RecipeBook by Shofo
@omri_ben_shoham1 Hey Omri, when we collect data we do so in a way that does not require logins, agreeing to TOS, creating fake accounts, etc. Our bar for collection methods is quite high and our approach is supported by public data case law. There are terms on RecipeBook's website that provide a bit more detail as well as a takedown form if that's helpful!
@omri_ben_shoham1 @braiden_dishman1 appreciate the direct answer, but I think it sidesteps the actual question. "we didn't need a login to get it" is a defense of how the data was collected, not a defense of what the buyer is allowed to do with it afterward. those are genuinely separate legal questions, and a takedown form only helps the original creator, it doesn't tell me, the buyer, whether the dataset I just bought is safe to train a commercial model on without getting pulled into someone else's lawsuit later. is there anything in the terms that actually indemnifies the buyer, or does the risk sit entirely on whoever purchases the hours
RecipeBook by Shofo
@omri_ben_shoham1 @galdayan Self serve purchases don't come with indemnity coverage. Downstream risk sits with the buyer. (Similar structure most data providers use) On larger deals we can negotiate custom contracts with different reps and discuss indemnity. If you're interested happy to have that conversation.
@omri_ben_shoham1 @braiden_dishman1 appreciate the straight answer, that's more honest than most vendors would give. but it points at an odd shape: the buyers who can afford to negotiate a custom contract with indemnity are the ones who could probably absorb the legal risk anyway, and the self-serve buyer, the solo dev or small team paying $15-480/hour, is the one least equipped to deal with a downstream claim and also the one stuck with zero coverage by default. is a self-serve indemnified tier ever on the roadmap, or is the plan to keep that protection sales-gated indefinitely
Searched for a tricky cooking motion I'd been hunting for and the re-rank after a few votes actually nailed it. The no-sales-call, pay-by-hour setup is refreshing.
RecipeBook by Shofo
@juliachild Glad to hear it worked!! Curious, what were you looking for?