# Sparse Thoughts > Research notes on foundation models, evaluation, and ML in biology by Gal Sapir. ## Posts - [the other risk in selling bio data](https://sparsethought.com/2026/08/04/the-other-risk-in-selling-bio-data/): on bio companies selling data to frontier labs, the risk of losing the original mission, and the limited commercial half-life of a static dataset. - [benchmarking is the new data activation](https://sparsethought.com/2026/07/03/benchmarking-as-data-activation/): benchmarks as a way to turn messy domain data into a measurable, optimizable substrate for models. - [what i've read lately](https://sparsethought.com/2026/06/24/what-ive-read-lately/): a loose reading list around agents, memory, benchmarks, and the hard parts of making models useful outside language and code. - [what did they actually measure?](https://sparsethought.com/2026/06/14/what-did-they-actually-measure/): a nature medicine paper says general llms beat specialized clinical tools: a close read of what its main benchmark actually measured, and what it left out. - [maps of context](https://sparsethought.com/2026/06/12/maps-of-context/): running PEEK's context maps on top of my second brain: a few hundred stable tokens of orientation can shape how the next hundred thousand get spent. - [long-horizon tasks](https://sparsethought.com/2026/05/21/long-horizon-tasks/): two modes of working with agents: synced mode does the expensive work of defining the chunk, delegate mode runs cheap on it, and the cost is (probably) the point. - [curation all the way down: on clinical AI benchmarks](https://sparsethought.com/2026/05/16/curation-all-the-way-down/): the curation regression, the openness trade-off, and what a substrate worth evaluating against would actually need: on Medmarks. - [small changes](https://sparsethought.com/2026/05/07/small-changes/): three small workflow shifts pointing in the same direction: portable context, cross-model delegation, and agent-first development. - [on the comparator in clinical AI](https://sparsethought.com/2026/05/03/science-paper/): A 2024 finding in 2026, a buried result the paper doesn't engage with, and the wrong comparator: on Brodeur et al.'s Science paper. - [a second brain, week two](https://sparsethought.com/2026/05/01/second-brain-week-two/): week two with a memory MCP as second brain: the background dread is lighter, reviews are cheaper but shallower, and why i'm not ready to let the system auto-fix itself. - [a week with a second brain](https://sparsethought.com/2026/04/23/second-brain-start/): notes from five days of running a memory MCP across Claude Code, desktop, mobile, and Codex: what's in there, what's already not working, and why the corpus is mostly corrections. - [maps, territory and LMs](https://sparsethought.com/2026/04/11/map-and-territory/): Borges' cartographers, Baudrillard's stages of simulation, and Polanyi's tacit knowledge: on the skill of reading AI-generated maps without losing touch with the territory. - [building tools as procrastination: a CLI for citations in Google docs](https://sparsethought.com/2026/03/13/cite-tool/): A CLI citation manager that resolves papers and inserts citations into Google Docs, plus some thoughts on building small tools as productive procrastination. - [what i've read in 2026 so far](https://sparsethought.com/2026/03/07/reading-q1-2026/): Books, essays, and papers from the first quarter of 2026: Dostoyevsky, Steinbeck, and some good writing on the internet. - [a small tool for diagrams](https://sparsethought.com/2026/02/18/drawio-with-claude/): Building a small Claude Code skill for generating editable draw.io diagrams — and why investing in narrow, single-purpose tools is surprisingly high-leverage. - [narrating your blog with local AI](https://sparsethought.com/2026/02/11/narrating-your-blog/): Open-source TTS crossed a quality threshold. Here's a tool that adds audio narration to Jekyll blogs using Kokoro-82M — runs locally on Apple Silicon, takes an evening to set up, costs nothing. - [a second opinion](https://sparsethought.com/2026/02/11/a-second-opinion/): Building a Claude Code skill that gets a second opinion from a different model family — and what the first real test revealed about what AI review can and can't catch. - [opus 4.6 and two small tools](https://sparsethought.com/2026/02/06/opus-4.6-other-stuff/): First impressions of Opus 4.6, and two small tools—an interview plugin and a markdown annotator—for staying engaged with your own work. - [cognitive offloading, exoskeletons, and remaining sentient](https://sparsethought.com/2026/02/03/offloading-cognition/): How to use AI coding tools without losing the skills and satisfaction that make programming worthwhile. - [how we actually evaluate agents (health)](https://sparsethought.com/2026/01/29/evaluating-agents-in-health/): Applying Anthropic's agent evals framework to health—what worked, what broke, and where general advice needs adaptation. - [the entertainment is instagram reels (and tiktoks)](https://sparsethought.com/2026/01/26/entertainment-instagram-reels/): A brief note on how Infinite Jest's 'Entertainment' prediction has arrived in the form of Instagram Reels. - [how will we know the model did a good job?](https://sparsethought.com/2026/01/23/what-is-a-good-fm/): Why we wrote evaluation criteria before code—lessons from publishing a CGM foundation model in Nature. - [data activation thoughts](https://sparsethought.com/2026/01/17/data_activation/): How to transform structured medical data into reasoning traces that improve LLM clinical performance—patient similarity and contrastive approaches. - [What's Sparse Thoughts?](https://sparsethought.com/2025/08/16/whats-sparse-thoughts/): Introduction to Sparse Thoughts—a low-friction space for collecting and reflecting on interesting content. ## Indexes - [Sitemap](https://sparsethought.com/sitemap.xml) - [RSS](https://sparsethought.com/atom.xml) ## About - [About](https://sparsethought.com/about/): About Gal Sapir