Running GPT-6 Astra and Fable 5.1 side-by-side and having each critique the other's output yields better results than either model alone.
My current view is that GPT 6 Astra is not meaningfully better than Fable 5.1 for my personal work, but that using both side-by-side is nonetheless very helpful and additive. I have been using GPT 6 Astra and Fable 5.1 a bunch over the past two days, largely for policy analysis, memo writing, and simpler software (e.g., making dashboards and forecasting models) that still nonetheless seems difficult conceptually. Across a variety of tasks I've done, it's been fairly random and hard to predict in advance which of the two models will end up being better at the task. For the tasks that are the most difficult conceptually, I've found that doing the project in both and then having each compare notes and critique each other has produced way better outputs than either alone. I think a reasonable person could conclude either model is the "best model" and it depends a lot on their subjective views and specific tasks.