
Most people ask which AI model is best and expect a name back. That's the wrong question. Here's the one we actually answer, and the bench we've built doing it.
Every founder I talk to eventually asks some version of the same question: which one should we actually be using. Claude or GPT, right now, this month, for real work. They want a name. For a while, I gave them one.
I don't anymore, because the question is malformed. It's the AI-era version of asking which programming language is best, and anyone who's actually shipped something knows that question doesn't survive contact with a real project. Ask ten operators which model is best and you'll get ten confident answers, each one true for about six weeks, because the honest answer was never a name.
It's a job description, and nobody was asking for one.
Part of why the wrong question persists is that the industry keeps selling the right one. Every launch comes with a new leaderboard, a new benchmark chart, a new reason to believe this is finally the model that settles it. It never settles it, because a single number was never going to describe something that has to be patient in one conversation and relentless in the next.
Nate B Jones put out a video a few months back that reframed this better than I'd managed on my own. His point: stop ranking models, start getting to know their families. The lab is the family. The specific model, Sol, Fable, whatever shipped this quarter, is a member of it. Rank Sol against Fable and you're ranking cousins from different households for no reason that survives the comparison.
That reframe changed how I talk about this with clients, and it's the reason I finally sat down to write out what our own bench actually looks like, instead of just telling people to try a few and see what sticks.
Nate's framing splits the field by what each lab actually bet on, not by which one benchmarks better this week. One family bet on reinforcement learning stacked onto existing lines. The trait running through every member of it: literal, relentless, keeps going until the task is actually done. Inside that family, Sol is built for long, technical knowledge work. Luna is the cheaper sibling, built for fast, efficient coding. Same instinct, different job inside the household.
The other family bet on scaling pretraining into bigger, more general models. The trait: patient with ambiguity, real taste, thinks before it acts. Fable is the member of that family I've actually put through the paces on real client work. Mythos shares the lineage.
Nate calls the family trait "big model smell." Some families have it, some don't, and that's a different question from which one is smarter. A relentless family member without much taste can still out-execute a patient one on a narrow, well-specified task. A patient family member with real taste will out-perform a literal one the moment the task gets ambiguous, and most of the tasks that actually matter in a business are ambiguous by the time they land on your desk.
Once you're thinking in families instead of leaderboards, the question changes shape. Not which model is best. Which family, and which member of it, fits the job sitting in front of you right now.
Here's what that looked like on an actual week, not a hypothetical.
We were a few days into pulling together requirements for a client build, and by midweek I had three completely different problems stacked on top of each other. A prospect had asked a pointed question on a call, and I needed a sourced, defensible answer before I wrote back, not a confident-sounding guess. Separately, we had close to two years of scattered client material, old transcripts, half-updated specs, a Slack thread nobody had ever cleaned up, that needed to become one coherent brief before a meeting the next morning. And underneath both of those, the actual build was still moving, and a piece of it needed a second set of eyes before it shipped.
None of those three problems got solved by the same tool, and none of them required me to sit down and rank anything first. I already knew which member of the bench handles a sourced answer, which one handles synthesis across a pile of messy material, and which one handles the second opinion on code. The routing took seconds because the knowing had already been done, in advance, the slow way, by using each tool on the same two or three kinds of problems over and over until the instinct was just there.
That's the actual skill. Not picking the smartest model.
Knowing your own bench well enough that routing stops being a decision and starts being a reflex.
Claude Code runs point on the build itself. It didn't win that spot by topping a benchmark chart. It earned it by showing up for the hardest, messiest parts of the work until I trusted it with the real stuff, which is a different kind of proof than a leaderboard number and, for my money, a more honest one.
Codex is the cousin. Same instinct for code, different lineage, and it's who I hand things to when I want a second opinion on what Claude Code already built. Not a fight between them. A second set of eyes that doesn't share the first set's blind spots.
Perplexity is still where I go when the answer needs a source attached to it. I don't trust any model's memory when I'm about to put a claim in front of a client, and I'd rather spend the extra ninety seconds getting it sourced than spend an hour later walking something back.
Gemini earns its seat for synthesis, not size. I used to describe this wrong. I said it was about the context window, how much it could hold at once. That's not actually why I reach for it. Point it at a pile of scattered material and it doesn't just skim the collection, it actually learns it. NotebookLM specifically, the part of that suite I live in for this, turns a stack of sources into something I can actually think with, not just search.
None of the four made the bench because they topped a leaderboard the week I picked them. They made it because they kept earning the next invite, on the same handful of problems, over and over, until not using them for that problem started to feel like a mistake.
There's a failure mode on the other side of this too, worth naming. Loyalty out of habit is its own trap. The bench isn't fixed because I like these four. It's fixed because I keep checking, and so far nothing new has out-earned an incumbent on its specific job. The day something does, the bench changes.
The families don't get my loyalty. The members do, and only for as long as they keep it.
Here's the part that actually matters for anyone running a business instead of just following the news. Leaderboard position resets every few weeks. Someone ships a new model, the chart reshuffles, and if you were optimizing for which one is smartest right now, you're back to square one on a schedule you don't control.
Knowing your own bench doesn't reset.
Every week you spend routing the right problem to the right tool, you get faster at the routing, and that speed doesn't go stale the next time a new model drops. It just makes the new model easier to evaluate, because you already know exactly what job you're auditioning it for.
If you're running a team and still find yourself asking which one you should standardize on, that's the tell. There's no single answer to standardize on, any more than there's one employee you'd hire for every role on a team. The standardization that actually helps is process, not tool: give your team the same permission I gave myself, to spend real time with each family member on the same kind of problem, more than once, before anyone decides anything. The bench builds itself once you stop looking for the one right answer and start collecting evidence on four or five candidates at the same time.
So the next time someone asks me which one they should use, I'm going to keep giving them the answer that actually helps.
Tell me the job first. I'll tell you which member of the family wants it.
Written by Skip Marshall
Learn more about our team