The Answers I Didn't Trust
Why I built a curated panel of experts to answer the health questions I wouldn't hand to a generic model.
Almost every “ask an AI expert” tool works the same way. Scrape the web for expertise, embed all of it, and trust that volume averages out to quality. On most topics that holds up well enough.
Health is where it breaks. The loudest content is usually the least rigorous: the SEO winner, the supplement affiliate, the person optimising for reach instead of for being right. Average all of that together and the confident tone survives while the rigor doesn’t, so ingesting more content makes the answer worse. The cost of being wrong here is a few years of healthspan rather than a bad paragraph.
I noticed it asking a chatbot what my VO2 max should be at 40. It gave me a clean number and a confident paragraph around it, and I was about to just believe it. Then I did the thing I should have done first and asked where the number came from. It came from nowhere I could point to.

Confidence without a source
A general model answers a longevity question in the same confident voice it uses to pick a pizza topping. It sounds sure, and it shows you nothing you can check. The VO2 max number I got may well have been right. I had no way to find out, and on a question about my own body that comes to the same thing as wrong.
An AI interface decides, without ever announcing it, whether its output reads as a suggestion or a command. Most of them default to command. I didn’t want a command. I wanted to know who was talking.
What it is
CortexPanel is a panel of experts you can ask real questions of. You pick a domain (health is the most built-out so far, with finance and psychology partway there) and instead of reaching for the open web, it answers from a corpus I curated by hand. Peter Attia, Andrew Huberman, Tim Ferriss, Rhonda Patrick, and a bench of medical specialists behind them. Roughly 65,000 chunks of their actual work, embedded and searchable, with every answer citing back to the source it came from. Ask what my VO2 max should be at 40 and the answer comes back framed the way Attia frames it, with a link to the episode or the page in Outlive where he says it.

One thing I learned building the retrieval: raw vector similarity is a liar. It scores how closely a passage matches the words in your question, and nothing else. So a nutritionist with one stray sentence about cholesterol would outrank the actual lipidology expert, whose relevant thinking is spread thin across thousands of chunks and never concentrated in one perfect-sounding passage. I had to build a blended ranking that weighs how central a topic is to a given expert, not just how well one passage matches. Getting the right person to answer turned out to be most of the work.
The curation
Here’s the part I’d defend hardest. Any decent model plus a vector database gets you a working demo in a night, and mine did, at 2 AM. What takes longer is the list of who’s on the panel and who isn’t, and that list is where the value sits.

Expertise gets curated in on purpose, filtered for integrity and signal density rather than popularity. That filter is a human judgment I can’t automate away. Some genuinely popular voices don’t make the cut because they sell more certainty than they’ve earned, and for each one I can say why.
No affiliate links
There’s a layer I’m still building called Picks: specific supplements, tools, and products for a given situation. It will never be tied to affiliate revenue or sponsorship. That decision costs me the obvious way this kind of thing makes money, and I’m keeping it anyway.

The moment a recommendation earns me a cut, curation has become advertising, and I built the whole thing to get away from exactly that. Once you know the recommendation pays the recommender, you can’t unknow it, and neither can anyone else.
Where it’s wrong
I don’t want to oversell it, so here’s the failure rate. I built a formal accuracy audit: 56 questions across five experts, scored by a separate model on grounding, retrieval quality, confidence calibration, and citation correctness. The baseline pass rate is 79%. Attia and Lex Fridman score 100%. Rhonda Patrick scores 30%, and the reason is boring and honest. She has about 1,200 chunks of content in the system against Attia’s 17,000, so the panel simply doesn’t have enough of her to retrieve. The most common failure across everyone is citation correctness: the answer is good and the citation lands in the right neighborhood, but not always on the exact claim.
Why publish the number? Because a health tool that hides its error rate is worse than no tool. If I’m going to ask people to trust curated answers over the confident average, the least I can do is publish where the curation runs out.
Why I keep working on it
It isn’t launched. It has one user, which is me, and the VO2 max question is one I genuinely asked it. The honest answers to the things I care about were scattered across a dozen people’s life’s work, and none of the tools promising to synthesize them were built to keep the sources honest. So I built the one I wanted to use.
It answers that VO2 max question now in Attia’s framing, with the page in Outlive where he says it. I still open the page.