LELenny's PodcastSep 24, 2026· 18:32

How to build products on a moving frontier | Dan Shipper (Every)

Dan Shipper, co-founder and CEO of Every, argues product orgs should add a small labs team exploring frontier AI like Astra 6 and Fable 5.1 while the product team executes, because exploring and executing pull in opposite directions. Labs teams of one or two people—messy "pirates" with "architects"—should dogfood, run parallel experiments, and expect to discard 90%, recouping ROI through content and early adopter programs. Anthropic Labs produced Claude Code and skills this way, and OpenAI's Codex, launched February 2026, merged into ChatGPT's 800 million daily users. Every's "Kate Bench" shows the pipeline: an AI copy editor trained on editor-in-chief Kate's edits cut her work 12% in a month. He advises weekly pipeline reviews and criteria—recurring usage, 10x better, affordable.

  1. 0:00Intro
  2. 2:05The core problem
  3. 2:44Why a lab
  4. 7:29Running the lab
  5. 10:56The pipeline
  6. 12:16Kate Bench
  7. 15:16Decision criteria
  8. 16:56Merging winners
  9. 18:01Closing

Powered by PodHood

Transcript

Intro0:00

Dan Shipper0:06

Okay, so I'm going to talk about how to build products on a moving frontier, um, or as I like to call it, the unreasonable ineffectiveness of business as usual during technology revolutions and what to do about it. So, the world just changed.

The world just changed, again. Uh, we had Fable 5.1 launch last week, we had Astra 6 launch last week, and these are fucking sick models. Like, look at this. This— I made this, like, Battle of Waterloo 3D, like, historical re- reproduction that's historically accurate, with, like, a single prompt and it just churned for, like, 4 hours, and then I got that.

Like, that's crazy. That is crazy. Someone else on my team made this, uh, sim— agent simulation where he fed it a scientific paper about how agents interact and, and how they get simulated together, and then he had, like, 1,000 agents all running on his computer and, like, visualized by this after, like, just a couple prompts and putting a paper in.

It's also now starting to do our video editing. We do a lot of videos, and, like, Astra is actually good enough to go into Premiere and, like, do a bunch of editing. And, uh, also, fun fact, the— all of the animations in this deck, and a lot of this deck was made by Astra.

Like, I did not touch a lot. I didn't touch the Animate tab at all, and we'll get into that in a sec. So this is crazy. What should you do? Do you keep your head down and focus if you're a product leader and you have a product team?

That might work for a little while. Uh, do you ask your customers what they want? Like, maybe, but most of your customers are not familiar with this stuff. They're actually looking to you, uh, to tell them what to do.

Do you turn your B2B SaaS app into a 3D multiplayer strategy game? Maybe. Maybe, but it's hard. This is— this is hard. This is a— this is a big problem that I think everyone in this room faces. I face it.

So the first thing to know, rule number 1, is never make any major life decisions. Within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a frontier model. Allright. We're settled. We're settled. That's rule number 1.

So let's get into the problem, though. So what's the core problem? Why is this so hard? It feels like, if you're running a product team, it feels like it's everyone's job to execute the roadmap at a high level and to stay at the frontier at the same time.

The core problem2:05

Dan Shipper2:19

And that's really, really hard because these are very opposed ways of working. Uh, if you want to explore the frontier, you're exploring. That's divergent. You're doing lots of different things. You're probably going to throw a lot of this stuff out.

You're— you're trying all these new models. You're doing demos. Exploiting is very different. You're converging. You are focusing. You're saying no to things. You're trying to execute against a roadmap that you've, like, you've planned out. And doing both of those, uh, you're— you're pulling in the opposite direction.

Um, so what should we do? Well, the thing I want to talk about in this talk is I— my contention that you should run your product org like a research lab, or you should add some research lab— lab elements to your product org.

Why a lab2:44

Dan Shipper2:59

And we're going to go through why you should make a lab, how to run a lab, and then what— what, like, how to bring what works from the lab back into your main product. We're going to use examples from stuff we do at Every and from some of the, uh, best companies in AI and how they do it.

So let's start. Why should you make a lab? The problem is, as we said before,right now, everyone in your product org is doing both exploration and execution. The yellow circles are the execution and exploration; it's purple, you can see everyone in your org is doing both.

And, uh, what— what's really interesting, though, is there are some people in your org, I guarantee it, there are some people in your org who are doing way more exploring than others. And that's— they're really important to identify.

Um, because if you identify these people, they are your early adopters. We are your early adopters. And by the way, that animation and this whole— all this stuff, all built by Astra, round of applause for Astra. Incredible. Incredible.

I did not do any of this. So we are your early adopters. Um, we love new things. We're already running Astra and Fable on the weekends to do our own personal projects or do things that are similar to, uh, to— to explore things that we might bring to work.

And because of that, there's a lot of things we might be excited about that aren't useful, but because of that, we're going to have a really good sense of what might— what might your product become, because we're, like, living in the future already.

The problem is that we can be a massive fucking distraction. And, um, the— so the— the key question to— to ask yourself if you're running a product org is, how do I harness the early adopters in my te— on my team, and maybe even, uh, some of my— some of my customers who are early adopters, without distracting everybody else?

And the solution that, uh, we found at Every and that I think a lot of big— uh, a lot of really great companies are starting to do is make a labs team. On a labs team, what you've started to do is separate concerns.

Some people are in charge of improving and scaling what already works. They're— they're on the product team. And some people are in charge of exploring what— what— what comes next, exploring the frontier, doing all these demos, trying these new models, all that kind of stuff.

That allows you to get the best of both worlds, and I'll— I'll explain how that works and why. And if you're looking at this and you're thinking, "Oh, that's really expensive, and I don't have the resources for that," and all that kind of stuff, what is amazing about AI is it allows you to have a labs team of one.

Uh, you can just have one person on your team whose job it is, maybe even, like,right when a new model comes out, to go explore all that stuff and come back and tell you what they learned. It actually does not require a lot of resources because all the people on your team are now so— they have so many superpowers because they can just go off and have Astra and Fable do a bunch of stuff that it doesn't require a huge investment to do this.

So again, to sort of pull apart the difference between a labs team and a product team, on a labs team, your job is you're going to explore capabilities, especially when new models drop, to see what's now possible. You're going to run a ton of experiments in parallel, and we'll get into that.

But crucially, the expectation for a labs team is very different. On a labs team, you are going to expect to dispose of, like, 90% of what you make. You try it and you throw it away. Product team's very, very different.

Your job on a product team is you're going to improve and scale the product. You want to deliver for existing customers. You want to make sure that they feel like they understand the product, it's coherent. You're not just throwing a bunch of garbage at them.

So for a product team, you want to expect to adopt about 10% of what the labs team makes or tries. And this is how they sort of start to work together. And there's a lot of really good examples of this starting to work in AI in a way that, like, we've had labs teams for a long time, but it's starting to work in a way that it has never, ever worked before because of this technology.

So a really good example is Anthropic Labs. Um, I'll be talking with, uh, some people at Anthropic a little bit later about this, but— and Anthropic Labs' Claude Code, like, one of the biggest, most successful productivity products of all time, came out— came out of Anthropic Labs, a separate little group in charge of doing lots of experiments.

MCPs, skills, Claude design, all of these things came from a small group of people inside of Anthropic experimenting with things, and then Anthropic investing more and more into the winners and the ones that worked. And there's, like, 1,000 experiments that you never— you've never seen.

Let's say you're convinced now that this sort of structure gets the best— gets you the best of both worlds. It— it helps you to unleash your, uh, your early adopters without distracting the people on the product team who need to actually deliver on the roadmap.

Running the lab7:29

Dan Shipper7:29

The question is, like, how do you run a lab well? So the first thing that I've started to come to is you really want to use very, very, very small teams. For a long time, the, uh, the standard for team size was the— the two-pizza team.

So that's, like, 8 to 10 people. That's what Jeff Bezos came to. I think in the AI age, I— I call it a two-slice team. Like, you want one or two people max. You can get so far, so fast with one or two people that anything more, there's a lot of co— coordination overhead and there's differing visions and it just doesn't work.

One or two people is really, really great. And I think the— the composition that works for us, and I— I think is starting to work for other people, is I— I call them pirates and architects. So one person on the team is the pirate.

They are slop cannons who are absolutely obsessed with— with finding value. That's— that's me. Like, I'm just out there just doing lots and lots of stuff and I'm throwing it away and it's all messy, but I'm going to find something really interesting.

And then architects are people who are like looking at a messy system, like maybe something that I built that's completely vibe-coded, and trying to help shape it into something that's valuable and beautiful and extensible. And pairing those types of people together, I think, is very, very powerful in AI.

Um, the next best practice that I think is really important is dogfooding. The biggest thing that— that matters on a labs team is making the feedback loop as tight as possible between making something and knowing if it's good.

And the tightest feedback loop is making something for yourself. If you can do that, I would do it. If you can't, then it's really important to get a couple of early customers or early adopter-type customers, uh, to help— help you test with a tight feedback loop so that you can iterate really, really fast.

Um, and to use the— the stuff you're building, all the experiments, ideally should be used for actual work that you actually have so you can tell, is this useful or is it just new? It's a big thing you need to differentiate.

Another best practice, which I think is quite different from how you should probably run a product team, is you should be building many experiments in parallel, even if those experiments are trying to do the same thing. Try competing approaches to the problem.

Uh, it— it might look like a mess and sort of, like, a lack of coherence, but everybody's going to have a little bit of a different take on how to solve a particular problem. And when capabilities move, the frontier becomes very unknown.

And so having people try the same thing from different perspectives helps you map that frontier and decide what's valuable and what's not. And then the last thing that I think is really important for doing labs is to figure out ways to make your experiments net positive in terms of ROI.

So we're going to throw out 90% of what we make. How do we make even that 90%, even the— even the 90% that doesn't make it into the product eventually, how do we make that actually useful for us?

So one thing that we've done a lot at Every that has worked really well is we turn them into external content. People love to see, "Here's all the things that— that we tried. Here's what worked. Here's what didn't."

And then that brings customers to— to use to read our stuff, use our products, all that kind of stuff. If that's possible, I highly recommend it. Another thing to do is use these experiments to feed an early adopter program.

There are probably a lot of, uh, customers that you have that want to be closer, want to be in the fold, who are early adopters, and this can be a value add for them to get to work with you more closely.

And then the last thing is, like, at base, the labs team should be sharing what they learn about capabilities and what's now possible with the product team so the product team can take that into account, uh, as they build things without getting distracted by having to go, you know, explore the frontier themselves.

The pipeline10:56

Dan Shipper10:56

So let's say you've done that. You're, uh, you've decided to make a lab. You're starting to figure out, uh, like, all the— all the little best practices. Now, how do you get the stuff that you're making into— from the lab into your product?

The thing you need to do is make a research pipeline where ideas start on the left and they start out as lab only, and there are lots of little experiments, and then you move them progressively from the left to theright, and at some point you have a handoff with your product team where they start to get incorporated into the product.

So to give you a sense of what this looks like, you've got the— the labs team. They're running lots of— lots and lots of experiments. Um, and most of them are not good, but a few of them might work.

And those become things that you start to test in real work. One of the things that we do all the time internally is we just let other people on the team adopt it and see if they like it.

Uh, there are different configurations that work depending on what kind of customers you serve, but the first step for us is, do other people on the team start to use it? If that's the case, it's often then ready to put in front of early customers and ready for the product teams to, like, really take a look at it and figure out, "Where does this fit in our roadmap?

And, um, how might we use this?" And, uh, and even there, not everything makes it, but, like, at— at some point, uh, the thing that you do will be ready to scale and release to customers, and we'll go through some examples of how that works.

That's the, uh, the research pipeline. And I'll give you, like, a little bit of a concrete example. This is Kate, our editor-in-chief, who I've been trying to automate for the last three years, um, very lovingly. Um, and Kate— Kate is fantastic at a lot of things, but one of the things she's really good at is she has a very good taste for copy edits.

Kate Bench12:16

Dan Shipper12:34

Um, but as the team has grown, we're about 30 people now, uh, she can't do all of the copy edits herself. Um, and we could hire someone, but it's really hard to get someone who has the level of taste that she has and train them and all that kind of stuff.

So what I've been trying to do is, like, figure out, how do we make— how do we make AI, like, help her with this so she can expand her impact in the org without having to spend additional time doing copy edits really late at night?

And so Kate often wakes up to messages like this from me saying, "Not important, but I downloaded every single one of your copy edits over the last three years and had Fable try to do a copy edit on the latest piece."

Um, and I've literally been doing stuff like this for a couple of years, and it just— it just started to work where the capabilities are good enough that it's— it's actually, like, uh, Kate ready for Kate to actually use.

So we move from— we— we call it Kate Bench. So we're starting to move from, "Okay, we've done a bunch of experiments. Now there's one that's, like, starting to work. Now let's push it out into the org and see if we can get some internal use."

So we make— we have a— we have an Every agent. It's a— it's a— it's a single agent for your entire company. It helps your company get AI built and use all of— all of, uh, AI workflows. agents.every.to.

But, uh, we use it internally, and what— what happens now is Kate is starting to be like, "Okay, at Every, do a Kate pass of this draft, and it'll go in and it will literally file suggested changes like she would based on her historical edits and then improve over time."

And this is something that I built just in my spare time that just— it just sort of works, but it's like, now we're starting to see that there's some value for the organization because she's starting to adopt it and it's starting to spread to other people.

And that's a really good sign. So when we get to this, um, internal use stage, this is the time where I would then bring in an architect, and I have one in particular who I work with on my team.

His name is Yannick. He's fantastic. Um, and Yannick is going to— Yannick goes and takes this, like, messy thing that kind of works and makes it really great. So now we have a whole dashboard of, "Okay, like, for each document, how many, uh, how many suggestions were— were accepted?"

And then how much work is remaining, uh, for Kate to do after, uh, after we go in and do it. So it checks what are the edits she does after the— the agent goes and does the edits. And you can see we're improving.

We did, uh, she did 12% less work, uh, on these types of edits than she did, uh, the month before. And his job is to make it, like, a real system that we can improve and compound, uh, in an ongoing way.

Now it's starting to work well enough that we're starting to think about, "Okay, how do we put this in front of early customers? Does this sort of process work for things beyond copy editing?" I think it does, but, uh, you know, we've sort of taken it from it's only an experiment in the lab to something that we're ready to be like, "Okay, I think we can give this to customers."

Decision criteria15:16

Dan Shipper15:16

So a few best practices for how to continually push things, uh, through your lab pipeline. One thing that we do that I think is very important is we review the pipeline regularly. We have— I have a tracker in Notion.

Every week on our all-hands, everyone talks about what's going on in our pipeline. And I think that's crucially important because what you want to be able to do is you want your product team, the rest of the product team who's not playing around with everything, to know, "Here's what we think is interesting.

Here's what's starting to move up the pipeline so that you can start to think about, 'Okay, if this does actually prove out, what are the implications for the product?'" without having to then go, like, scramble to, like, do your own experiments and all that kind of stuff.

It helps keep, uh, the lab and the pipe and the product, uh, orgs in sync. Another really important thing is defining clear decision criteria for moving an idea through the pipeline and really making it an event when you do it.

So some of the decision criteria, like our big one, is, "Are people using it? And are they coming back?" Because what you're trying to— what you're trying to do is use, uh, internal use as a proxy for value.

The second big question, "Is it 10x better than what currently exists?" And this is actually a really interesting one, a really interesting filter because what is new feels very exciting, but the question is, when it's a month later, is it actually any better?

It's hard to know unless you use time and usage over time as a filter. Um, and then also, "Is it affordable?" Like, "Maybe this works now, but is it— can we actually serve it at scale for our customers?"

Another really big, big question. And— and having that clear criteria allows you to make sure that you're— you're being rigorous about how you do this. But the ultimate goal is that you, um, merge the winners into the main product.

Merging winners16:56

Dan Shipper16:56

And this is— this is this full cycle of labs to merging winners is something that is starting to happen really, really well. Like, we've had labs for a long time, but the idea of doing a self-destruction of building something that then disrupts your own product has been around for a while, but it is actually starting to work now in AI because you can move so fast.

I think a really good example is Codex. So, um, Codex built by a small team working outside of the main app. There were a bunch of different teams inside of OpenAI who were working on the future of coding at the same time as Codex was.

They had a— they did a bunch of different iterations on different form factors for what the future of coding would look like. Is it inside an IDE? Is it a CLI? All that kind of stuff. Um, but they launched their desktop app, which, uh, I think is fantastic.

They launched it in February 2026, and it grew so fast that they eventually just merged it into ChatGPT, and it became the foundation of ChatGPT. Like, they— the— the small little Codex team then just, uh, took over this 800 million daily active user app because they were so successful with this— this way of experimenting.

Closing18:01

Dan Shipper18:01

So I think that's the promise of doing something like this is that you can actually build the next version of your product while scaling the one that you already have, and you can do it without going crazy. The way that you know this is working is you will welcome moments when new models drop and be excited about them instead of dreading them.

And that is my talk. Uh, I'm Dan Shipper. I'm the co-founder and CEO of Every. You should check out Every. We're the only subscription you need to stay at the edge of AI. Thank you very much.