Journal 2026-2027

9.15.26

Dear Reader,

All labs are now fully operational and producing. I ran the first non sterile pre testing GPT-6 Astra evaluation through the revamped benchmark framework this week and the results were, to put it mildly, striking. Not only is the gap between Astra and its predecessors real and measurable, but the new framework captured nuances that Chop-Bench's original methodology would have missed entirely. Austin's work on this has been outstanding and I am very glad he chose this direction. The areas where Astra excels most are also the areas most relevant to our scraper research, which is worth paying attention to. Dash's VLM pilot continues at a steady pace, they have finished the initial data pipeline and are preparing for a first small-scale training run. It is modest in scope, as it should be, but the fact that we are doing it at all feels like a meaningful step for the group. Now, today's big news, and it is big. TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released a model called Jev. What makes it remarkable is that it is not an LLM. It does not generate text at all. Jev is what TypeSafe calls a "System One Model." You give it state and typed questions, and it returns structured decisions with calibrated confidence scores in a single parallel pass. It mathematically cannot hallucinate because it never generates strings. So Deterministic in content not structure SO DOPE, Anyway It is named after William Stanley Jevons of Jevons paradox fame. Apparently it can play Doom EEEP! The reason I bring this up beyond my obvious enthusiasm is that this has direct relevance to the game-based benchmark testing Lucian and I have been piloting. A model that returns typed probabilistic decisions rather than natural language is, in some ways, exactly what you want for game-based evaluation environments. I will be watching Jev closely. On the INSIGHT front we had a productive conversation about the viability of a new study design for the coming semester. Google is also due for some kind of response to Astra. The model release pace this year has been relentless and, as I have said before, this proliferation is both a blessing and a challenge for a research program like ours. We have to be disciplined about what we investigate and what we set aside. I am proud of the team. Austin, Dash, and Lucian have all met the challenge head on.

Per aspera ad astra,

All the best,

Guy Gendell


9.10.26

Busy week. Austin has made excellent progress on the Benchmark Information framework redesign. He's been restructuring how we define evaluation criteria so that when a model like Astra drops, we have a standardized protocol rather than scrambling to set up ad hoc tests. The first Astra evaluations through the improved framework should be ready to run within the week. I am very pleased with the direction he is heading. Dash's VLM pilot is moving along, they've been working on data preparation and preprocessing which is, as anyone who has done ML work knows, the least glamorous but most important part. Dash is learning a lot about the practical realities of training pipelines and I think that experience alone justifies the project regardless of what the pilot model actually achieves. In other news, I've been doing some pilot testing with Lucian on game-based benchmarks. The idea is deceptively simple: use structured games as evaluation environments for model decision-making. Games have clearly defined rules, measurable outcomes, and bounded state spaces, which makes them surprisingly good test beds. We ran some very early trials and while it is far too soon to draw conclusions, the initial results are interesting enough that I want to explore this further. It's a bit of a side project for now but could tie into Austin's broader benchmark framework work down the line. I have also been iterating on SIGNAL's gesture recognition pipeline, the latency between Mediapipe hand tracking and the command normalization layer could be better. We are exploring whether a lighter model through Ollama might reduce that lag without sacrificing accuracy.

All the best,

Guy

9.5.26

Dear Reader,

And there it is. GPT-6 Astra dropped on September 3rd. OpenAI is calling it their "most intelligent and aligned model" and their president went so far as to say we have entered the "AGI era." I have thoughts. First, the model is genuinely impressive on benchmarks: 99.9% on ARC-AGI-3 (oy Vey), near-saturation on FrontierMath, and 100% on ExploitBench. Second, the cybersecurity angle is extremely relevant to us. Astra is the first model OpenAI has classified at the "Critical" cybersecurity capability level, and the full version is gated behind a trusted access program called Daybreak. The public model refuses advanced cyber tasks. This is notable. Third, and this is my hot take, calling this AGI is premature at best and marketing at worst, however impressive the benchmarks may be. But I will leave that debate to people more qualified than a high school researcher. What I will say is that this is exactly the kind of release our benchmark framework needs to be ready for. We are already working on integrating Astra into our pre-testing pipeline backlog . Austin's timing on this initiative could not have been better. This is going to be a big semester.

All the best,

Guy Gendell

8.29.26

We've been getting organized across all labs this week and it feels good to have some structure back. Austin has formally taken on the benchmark framework and pre testing relevant intelligence improvement initiative. His first goal is to audit our existing Chop-Bench methodology and identify where the framework can be made more rigorous and more generalizable. The idea is that we should be able to spin up a structured evaluation of any new model within days of release, not weeks. I think this is incredibly valuable work, especially given what we saw this summer with the pace of releases. On the Dash front, we've started scoping the VLM pilot. Dash has been doing some background reading on training pipelines and data preparation, and we are looking into what compute resources are realistically available to us. This is genuinely new territory for WISRD and I want us to approach it carefully. Ivan and I checked in briefly about SIGNAL, the Mediapipe gesture recognition still works well but we need to figure out our goals for this year before we start building again. I also met briefly with some new WISRD associates, several of whom expressed interest in the AI and CS space. We shall see who sticks around.

Best,

Guy

8.26.26

Dear Reader,

First full week in the lab. It is always a bit chaotic getting everyone oriented and reacquainted with where things stand. Had our first round of lab meetings with Gen AI and INSIGHT. Austin and I had a particularly good conversation about where he wants to focus this year. With Leo's departure we have had to rethink how we allocate our efforts, and Austin has expressed a strong interest in improving our benchmarking frameworks, not just Chop-Bench specifically but our methodology for evaluating models more broadly. I think this is the right call, especially given the pace of model releases this year. Dash meanwhile has been pitching me on a pilot project involving Vision Language Model training. I was initially skeptical given it is somewhat outside our usual scope, but the more we talked the more I saw the value. VLMs are increasingly relevant to how AI systems interpret and interact with visual data, and a pilot training project could give us firsthand understanding of the pipeline rather than just evaluating the outputs. We will discuss more next week. On the topic of browsers, because I simply cannot help myself, Cloudflare released Kitesurf earlier this month. A browser built entirely for AI agents, not humans. No tabs, no extensions, no pixel-perfect rendering, just machine-readable content extraction running on Workers. Uses reportedly 3-7x less CPU and memory than Chromium. Significant implications for the web scraper project.

All the best,

Guy

8.22.26

Dear Reader,

Welcome back! The 2026-2027 year begins and I could not be more excited. This summer was productive in ways both expected and not. Chop-Bench ran its first full round of funded testing over the break and the data, while still being processed, looks promising. The CLI is still annoying, but the local UI I caved and built in April has made things significantly more manageable. SIGNAL is in a good place architecturally after the Mediapipe integration and model upgrades from the spring, though Ivan and I will need to sit down and discuss what his goals are for it this year. The web scraper project's transition to TypeScript, which felt so painful at the time, has proven to be the right call. On a less positive note, Leo D. will not be returning to the Generative AI group this year. I want to thank Leo for his contributions last year, We wish him all the best. That said, Austin is back and more motivated than ever, and Dash has been talking to me all summer about some very interesting ideas. More to come.

All the best,

Guy Gendell