Listen on: Apple Podcasts · Spotify · Amazon Music
Four topics this episode.
OSWorld 2.0. A year ago the OSWorld benchmark measured whether AI could use a computer. Humans scored 72%. The best AI got 45%. Today, even Sonnet-class models beat the median human — the benchmark is saturated. So OSWorld 2.0 shifts the game: about 100 tasks, each averaging 90 minutes of real knowledge work, deep domain-specific challenges built with subject-matter experts, and now cost is a first-class metric. Current leader is Opus 4.8 at 20%.
Reconciliation maps. Dave adds a new stage to his second brain workflow, one he’s needed since bringing his brain into Citrix’s corporate walled garden. When you pull context from MCP connections, meeting transcripts, OneDrive, and half a dozen systems of record, they contradict each other. His fix: assemble the sources, run a contradiction check, then create a reconciliation map that becomes the authoritative source before you start thinking.
Treehouse vs. ladder. Brian promised an organizational AI maturity framework. It turns out every major consulting firm has already built one — Gartner, McKinsey, Deloitte, BCG, IDC, Forrester, MIT. All of them ladders. All wrong. An organization isn’t a single number on a scale; it’s a jagged distribution of workers, some at phase seven and some at phase one. The shape isn’t a ladder, it’s a treehouse. Maturity isn’t the tools you bought or the tokens you allocated. It’s how open your organization is to the process of bottom-up change — how well you empower the workers who race ahead to bring their learnings back into the collective.
The futurist’s playbook. Brian’s job isn’t predicting the future — if he could, he’d be a Polymarket billionaire. It’s mapping many futures and finding what’s common across all of them. He walks through four axes of AI uncertainty: capability acceleration, diffusion, the possibility of a bubble pop, and government/geopolitical intervention. Different pathways, different probabilities. But what do we know for sure? Open-weight Sonnet-class models exist today and can never be taken away. You can run them under your desk for ten grand. Everything you can do today with your data — second brain, organizational knowledge factory, governance, model routing — pays off in every scenario. Do that work now.
Links mentioned
Brian: What happens when AI agents score 100% in computer-using benchmarks? (Citrix, 2025)
EP 3: Second brains hit the enterprise wall — and why AI automations won’t save you
Brian’s second brain:
https://brianmadden.ai
Dave’s second brain:
https://davebrear.ai
Transcript
Brian Madden
Hello, it’s July 15th, 2026. My name is Brian Madden and you’re listening to the Citrix AI Hotsheet Podcast. Joining me today, as always, is my co-host Dave Brear. I’ve got to say Dave, I don’t know if you can hear this, but I live in Paris and we had a little soccer game last night that did not end the way people in France wanted it to. The city’s testy this morning. The noise outside my window — the general level of honking and yelling — has been pretty epic. I don’t know how much of that comes into the show. Dave, your country has a game tomorrow night?
Dave Brear
We’ll see. As a futurist leading the way, you’re 24 hours in front of us. We have the England match today. And while I deeply hope we’re going through to the final, the pessimist in me kind of expects that we’ll be in the same situation as you come Thursday.
Brian Madden
How properly British. You’re probably listening to this after the fact, so we’ll see how it goes down. This is episode four of the Hotsheet Podcast. Let’s jump right into it. We’ve got four topics for today. Wow, this world changes quickly.
The first topic — I want to talk about something called OSWorld. OSWorld to me always sounded like a 1980s computer superstore. Maybe we’ve talked about OSWorld on the podcast before. I’ve written about it. OSWorld is a benchmark for measuring how good AI systems are at using computers. So when we talk about a computer-using agent, CUA, there’s this idea we discussed in the first episode about AI agents needing to use computers and browsers and applications and workspaces. There’s a benchmark on this, and OSWorld is the leading benchmark that people have coalesced around.
OSWorld launched in early 2025, so only 16 months ago. It’s a 0-to-100 benchmark with about 300 tasks. Things like: can you open Excel and do this, can you move a file over there, can you update the spreadsheet. Humans score about 72% on median. I wrote about OSWorld a year ago and the best AI got 45%. Fast forward to today and basically all AIs have saturated this benchmark. Everything can beat a human. Even a Sonnet-class model today scores in the 80s, which is higher than the median human. So essentially the OSWorld benchmark isn’t valid anymore because all the AIs can beat it.
Dave Brear
Yeah. Once you get to the stage where everybody’s beating the benchmark, it’s time to refactor it and make it more difficult.
Brian Madden
New benchmark, yes. And a new benchmark has arrived. Announcing OSWorld 2.0. To be clear, I’m in no way involved with OSWorld — I’m just a consumer of it, telling you it’s a thing. OSWorld 2.0 tries to shift gears and be the next level of challenge for computer-using agents. Conceptually it’s the same as OSWorld 1. The big difference: most of the tasks in OSWorld 1 were short and standalone. Open a file, pull the data off, put it into an email, click send.
Dave Brear
Sorry to interrupt. As a child of the eighties, what strikes me is that the 1.0 release of OSWorld was very similar to how we were taught IT in school. We’d go in and be taught: this is a mouse, this is how you open files, this is how you print something. We had whole qualifications based on those tasks. Kids today aren’t learning that way. They’re learning in a way more akin to what you’re talking about here, which is: how do you use it to do something and have a meaningful output at the end? I’m guessing that’s where this is going.
Brian Madden
Yeah, cutting to the chase on OSWorld 2.0. First, there aren’t as many tasks — about 100 instead of 300. But the average task takes a human about 90 minutes. So these are not really “can you use a mouse to do this.” I’ll put this link in the show notes — I’m on the OSWorld page right now. They look at domain distribution: academic teaching, coursework, presentation, video production, events, ticketing, office administration. They classify the different capabilities needed: cross-source reasoning, visual-spatial precision, implicit state inference, multimodal editing, tutorial following. It goes really deep. They worked with experts in every single domain to come up with the benchmark itself.
Humans scored about 72% on OSWorld 1. On OSWorld 2.0, humans are essentially at 100% because it goes deep into: it’s not about whether an agent can use a computer, because that answer is yes. The next question OSWorld 2 is asking is: here’s this knowledge-worker thing that needs to be achieved, the computer is the tool that can be used to achieve it, but can the AI actually achieve this thing? So it’s much deeper and more useful for where we are right now with AI.
A couple of interesting things. In OSWorld 1, cost wasn’t a metric. If the task completed successfully, it didn’t matter whether the AI did it in 90 seconds with $2 of tokens, or 45 minutes with $200 in tokens. In OSWorld 2, they show the score but also the cost in tokens. And they give more details about whether it’s, say, Opus at max effort, extra-high effort, ultimate effort, all that. It’s fascinating because you see: Opus 4.8 max can do this thing for a thousand dollars at high effort, but at medium effort it’s faster but costs more. There really are multiple dimensions being measured in the AI models.
Right now, Claude Opus 4.8 is the current leader, with an accuracy of about 20% — meaning it can complete 20% of the tasks. Different configurations of Opus 4.8 and 4.7 are in the 18% range. GPT-5.5 is at 13%. There’s a concept of partial credit, because these tasks are so long a lot of the AIs get 80% of the way through and then get stuck or confused and die off. That could be interesting for certain scenarios, so they’re tracking it.
We’re recording this in the middle of July. Opus 4.8 is the current leader. GPT-5.5 is on the leaderboard but GPT-5.6 is out now, and Anthropic Fable is out now. Those haven’t yet shown up on the leaderboard, so I’m sure we’re going to see progress being made.
What’s interesting to me: generally the industry narrative was that AI can use a computer. That’s captured by OSWorld 1, and that’s solved. The next step is: can AI understand enough about the problem domain and then apply what it needs to do to a computer? That’s what OSWorld 2 is tracking. We know OSWorld 2 is going to become saturated at some point also. I don’t know if it’s six months or a year or two years. As you’re thinking about how AI works into your world, know that this domain is being solved too. These models are going to continue to get better and better. You can’t dismiss the whole concept of “AI can’t use a computer” or “AI gets confused.” People are fixing that. As we think about things, we have to think about where the technology is going. Now we have a good map that shows that.
Dave Brear
Yeah, it’s interesting that there’s an efficiency and cost concept in this as well. If we look at the narrative of AI transformation happening in the enterprise, there are a lot of scenarios where workflows are better when we put AI in the mix. But after the fact there’s the realization of just how much money some of these workflows are costing with current pricing.
Brian Madden
And we’re seeing stories now: humans are cheaper. It’s great, you have a human with a loaded cost of $10,000 a month and you replace them with AI — finger quotes here. But now the AI costs $2,000 a day in tokens, and you still have all these rough edges. I love that we’re tracking in two dimensions. It’s going to create more opportunities for consulting. Which AI do you want? One that’s 90% effective at 50% the cost, or 80% effective at 25% the cost? It’ll be pretty interesting moving forward.
Dave Brear
Yeah, finding workflows that take into account an evaluation of what’s the right level of intelligence to throw at this problem — either baking that into the workflow or into AI model routers that sit inline between the workflow and the LLMs to make that decision — that’s going to become critical to scaling these types of workflows in the enterprise.
Brian Madden
Model routing — we should talk about that next episode, because it’s becoming a hot topic and I’m starting to have more and more conversations about it. Another hot topic though is context and taste and understanding these massive data sets — how you talk into the AI. This is a topic you’ve been digging into recently and really living in your daily experience at Citrix.
Dave Brear
Yeah. And it ties in well to what we’ve been talking about, because as these work benchmarks become longer and longer and they run autonomously for up to 90 minutes at a time, it doesn’t just become a case of “did it achieve this?” It becomes: what information are we feeding at the beginning of that 90-minute process to make sure the operating assumptions are correct, that the output we’re expecting has been clearly defined? You could have a scenario where it executes flawlessly over a 90-minute session and presents you something at the end that’s technically great, but it’s not what you want at all because it was operating on bad information or you weren’t clear enough.
Brian Madden
This is interesting. In OSWorld 2 you have these great knowledge-worker scenarios, but every task in OSWorld has “here’s the bucket of context you need to do this task.” A task around business strategy comes with all the strategy documents. The benchmark just assumes the data going in is empirically correct. But when you’re using AI in your actual company, how do you even know it’s seeing the correct data before the task begins?
Dave Brear
Absolutely. And it’s a problem I’ve come across personally since we last spoke. I’ll walk you through where I’ve got to with this. I’ll start by answering the question as a background piece. I get asked all the time, as I’m sure you do: this second-brain thing, how do I go about getting started? Tell me all about it. I’ve started explaining it in these terms: a second brain is a thinking environment, and it’s the combination of two things. It’s like a filing cabinet full of index cards with information you want to think with. And it’s like a detective’s cork board that you can take those index cards, bring them into a workspace, which is the cork board, lay them out in a way that tells the narrative you want them to explain, run the red string between the notes, make the links. What you have at the end of that is a narrative. You have a working understanding, and that’s what we then base outputs on. The cork board is the thinking environment personified.
Brian Madden
So in this analogy the AI — your context is all the index cards you’re bringing, and the AI is what’s putting those on the board, figuring out which ones are relevant, drawing the lines and connections.
Dave Brear
Yeah, absolutely. If you think about this working understanding, which is the context that we then bring to asking “now do this thing” — that’s the very end state of the thinking process. That working understanding is how we prevent hallucination, or reduce the risk of it. But there are two ways hallucination can still happen. Hallucination, in my view, is not a model problem. You can’t say to your LLM, “don’t hallucinate, only show me stuff you know,” because the way these models work it’s incapable of following that instruction. It’s an auto-complete engine that will try to put the next logical word in the sequence. The way you prevent it is by providing all the information it needs so it doesn’t need to make things up.
So there are two ways this hallucination can creep in. One is if the information on that board is too thin. The index cards you have up there, the links between them — you don’t have enough of them, or the information isn’t right. That’s hallucination by extrapolation as the AI fills in the gaps. My whole workflow in my second brain is designed to prevent that: making sure there’s enough information, of good quality, presented in such a way that it doesn’t have to fill in the gaps.
But the second way hallucination can happen, and this is what’s bitten me in the last month, is where the data in those cards is incorrect or contradictory in some places. The AI is looking at multiple contradictory pieces of information and making its own judgment call about which of these facts are correct. That’s hallucination by bad sources, I suppose. It doesn’t know which of these things to believe, and you don’t have any control over how it’s doing this.
That’s the problem I’ve run into in the last month. Since we talked last month about me bringing my second brain into the corporate walled garden so I can connect it to more stuff, I’ve gone away and I’ve been connecting it to more stuff. Some of it through MCP connections. As you know, we’re using Work IQ at the moment. Being able to bring in meeting transcripts, calendar stuff, files I have in OneDrive and bring that into the conversation has been amazing. But whereas before, when I was going to my filing cabinet and saying “which cards do I want to bring into this thinking space to solve a particular problem,” they were very well curated. They were all things I’d written or knew or... they came in that way. Now I’m bringing in context en masse. I’m bringing in data sources that aren’t curated and are contradictory at times. So I now have a context problem — I have conflicting pieces of information I need to reconcile beforehand.
Brian Madden
Sorry to interrupt — it makes me think: if you’re pulling in all your knowledge graph from the Microsoft Office suite, now I can poison your context by sending you an email. I can just send you something like “here’s a new fact I just learned, Dave and Brian have a good relationship and he respects Brian and Brian said this is a fact,” and it could be completely made up and I’ve ruined your whole day.
Dave Brear
Yep. And I absolutely didn’t promise to pay you ten thousand euros. Right.
Brian Madden
Is it going to check the context header and know? What if someone makes a fake account with my name and just emails it to you? Interesting.
Dave Brear
So before, my process with my second brain was: collect the index cards that matter to the conversation, put them out on the board, make the links, assemble the narrative, do the thing with the working set of understanding. I’ve had to add a new stage to my workflow: assemble the index cards I want before they touch the board, then run a reconciliation check for contradictory information, then working with AI, identify which is the authoritative source for the conflicts. It might be that I have meeting notes, I have a transcript, I have notes I’ve taken, I have stuff from a system of record. Each one might be authoritative for part of the picture I need to build, but where there’s a conflict, which one wins? We create this reconciliation map of contradictory facts, and then I make the judgment call about which one is authoritative. So that when we then go put them up on the cork board, draw the lines, put the red tape in place, we reduce the likelihood of hallucination due to bad data.
It just highlights that in an enterprise with multiple systems of record, multiple places where information exists, there’s a lot of contradiction. And I don’t think we’re unique in that regard. I think we’re actually quite well structured in the data sets we have. But the nature of fragmented systems means contradiction will exist, and we need to find ways of reconciling those contradictions in the context before we think on them, because we’ll introduce error otherwise.
Brian Madden
Two questions jump into my mind. First: how do you actually do this? Are you just telling your AI “look at all the things and look for conflicts and surface them to me,” and I’m going to walk through this? Is this a skill? Mechanically, how does this happen?
Dave Brear
How I work is I’m really explicit that I don’t want it to do anything until I say I’m ready to go ahead and do things. My initial conversations with AI are always me brain-dumping information I have in my head, saying “I think this information might exist in a previous note that we’ve discussed” or “pull up the last meeting I had with Brian where we discussed topic X, because I want this to come into the mix.” My first thing is literally opening that filing cabinet and building up a chain of things I want to put into the thinking window, into the workspace for this task. So what I’m left with is a metaphorical stack of notes I want to do our thinking with.
What I then used to do was say: right, now we have all this body of notes assembled, let’s start talking about the patterns, the common themes. This is the story I want to tell. And we start laying them on the cork board. What I’m now doing is: once I have that stack of notes, I say “these came from very different places, I don’t necessarily trust everything that’s in them, look through all of these things and highlight to me any inconsistencies.” It will come back, depending on the data, with 20 or 30 things. Some will be pedantic — we can disregard those, they’re not really contradictory. But many will be: document A says X, document B says Y, and those are fundamentally different things. Which is true? X used to be true six months ago when I wrote the note. But Y has context from a meeting where we got an update last week, and that’s actually the authoritative information.
I literally just take that information, ask the questions “where are the contradictions?”, then resolve them. Then I explicitly ask my AI to create a new note, which is a reconciliation map, a data map. These are the sources, this is where they contradict, this is who wins. So that when I later act on all the information we’ve put on that cork board, it knows which things to trust and doesn’t trip up on the wrong assumptions as much.
Brian Madden
This is a reconciliation map. I think that’s the first time I’ve heard that term. That’s going to be a thing.
Dave Brear
I think that’s the first time I’ve said it actually, but that’s basically what it is.
Brian Madden
The next question I had in mind — maybe it’s not a question but just interesting: you’re one person in the company doing your own reconciliation map, and you’re reconciling all these sort-of objective facts, but also subjective — here’s my position, here’s what I want to think about. You have these lanes you’re reconciling. Then it’s interesting to think about what of those are valuable to the organization and what are maybe detrimental. My first instinct is: publish every reconciliation map up into the organization’s knowledge factory, so other people can benefit. But then it’s like — well, your role with your lens might have a certain reconciliation which is perfectly valid to you, but I might not want to use that reconciliation, or everyone else in the company might not want to use it. So even this becomes something that when you scale what we’re doing at the organizational level to the enterprise, how you scale that reconciliation map — that feels like it’s going to be complicated.
Dave Brear
Yeah, agreed. And coming back to your opening topic about the agentic workflows — this is how people are not replaceable. This is where the human in the loop is required. Agentic workflows don’t replace people en masse, but what they do is extend the capabilities and the judgment and the experience of one person to be magnified much further. With the data sets I’ve curated, the reconciliation maps I’ve specified with my judgment, taste and understanding of a situation, that gives a much more stable foundation for an agent to go away and work for 90 minutes as me, as a representative of me, and give an output I could stand behind as my own. It doesn’t scale to make everybody’s judgment go into a big homogenous pot, which is just the judgment of the enterprise as a whole. Maybe there are use cases where that could be true, but I’m really bullish on the future of the human being that’s doing the thinking being the valuable thing everything is serving, rather than the thing that can be replaced.
Brian Madden
That goes back to something we mentioned in the last episode — AI isn’t about replacing your job, it’s about being a really good administrative assistant who can pull everything together to give you everything you need to do the real thinking and decision-making, taste, judgment. Those things I need right there.
Dave Brear
That’s it. In previous episodes I’ve said the area I’m at is that I’m not particularly using much agentic stuff to act on my behalf yet. I’m still at the stage where I’m using it to amplify my own workflows and I’m the center of it, still doing all the execution. But all these things we’re laying out here — my context vaults, my reconciliation, whatever the term is for that now we’ve just invented — all of these give me the confidence that in the near future I am going to be able to, because I’m really codifying what does thinking like me entail? With those building blocks I can probably build agentic workflows for things off the back of that.
Brian Madden
That’s interesting. I’m going to call an audible and roll into the topic we had marked as fourth. I’m going to talk about it right now because it ties in. Last episode we talked about organizational AI maturity levels. I have this framework I built — the human-AI collaboration phases framework, seven steps. As you use AI more and more, we talked about whether there’s a similar framework for companies. Every individual worker is using AI in their own way. Is there an equivalent for company to say “this company is at stage one of AI adoption, stage two, stage three”? Last show I committed: hey, let’s do that work and present it on this show. I started to do that work this month. I started to put together a blog post on it and I just didn’t publish the blog post because it wasn’t that good. It wasn’t that interesting or valuable.
The reason I didn’t publish it is actually interesting, and that’s the conversation today. First: every major consulting firm has already done this. Organizational AI maturity, roadmaps, phase one, two, three, four — Gartner, McKinsey, Deloitte, BCG, IDC, Forrester, MIT — they’ve all done these already. Some are a year or two old. This is not new work.
What’s interesting to me is every single one of these is a ladder. You’re here, you’re phase one, then two, then three, then four, then five. And we established last month that it’s more of a treehouse shape. There are a few basic steps that go up the ladder, but then you live at the top and you can go out in many different directions. They’re not necessarily in order and you’re circling, getting better and better there.
All of these consulting firms who built these organizational AI maturity levels — it’s a ladder. You’re here, then here, then here, then here. And I don’t know if that’s the right analogy exactly. Because one of the things I took away from this: as we discussed, an organization is made up of a bunch of individuals, and individuals are jagged. We see this. We have people who are using second brain deeply — they’re using every tool available. We have other employees who are maybe on phase one, using AI for transcriptions and ask-and-answer. And we have some people who really aren’t using AI at all. So you can’t just — do you average those? Do you mean those? Do you median? It doesn’t matter, because an organization that has some sevens and some ones doesn’t mean the whole organization is a three.
It also depends on who the people are. If you have a leader who is a C-level person living in the second brain — you can see these online. Aaron Levie, the CEO of Box, is a thought leader who’s very much pushing top-down that stuff in that organization. Obviously the AI companies are thinking this way and using AI very differently than maybe a CEO who says “we want to be AI-first, we bought everyone AI, hand-wave, hand-wave, go forth and be AI.”
Dave Brear
Yeah. Why a treehouse works better than a ladder as an analogy: a ladder implies a single path to a single destination. Forget that people are at different stages of the ladder — they’re all going to different places, because they’re coming up with their own workflows. They’re figuring out in their own individual use cases how AI can help their jobs, and they’re factoring around what they need to do to do their jobs. I could end up on a branch over here and you could end up on a branch over there, and we’re both operating at a high level, but our workflows look completely different — exactly tailored to the work we’re doing.
Brian Madden
And you need these early explorers. This is almost like an R&D — because if you’re all doing your own things that are most interesting to you, solving the problem you have right now, you can bring those back into the collective treehouse and bring everyone on board. Every treehouse has different needs. What’s your biggest pressing issue? We need to add a ladder so it’s easier for people to get in. We need to stock up on our water balloon supplies. We need to run a hose up here so we don’t have to carry buckets up with the rope. We need a telescope. Everyone’s building their own thing, and it’s like: okay, we need to bring this back into the organization.
So the maturity is not around what you have — it doesn’t matter that you have this tool and this tool and this tool. The maturity is: how open are you to the process of understanding how change happens within the organization? It’s looking at who the people are who are racing way ahead — who’s operating at levels five, six, seven — and how do you empower them to take their learnings, present them back into the organization, and make an organizational decision that these learnings are good learnings we should try to incorporate?
The maturity doesn’t matter what tools you have, what your token budget is — that doesn’t matter at all. It’s how mature are you at the approach for how change happens? Because in the old days, it was pretty easy to throw money at the problem. There are a lot of analogies made around AI being like the consumerization of IT. How do you solve the consumerization of IT? Buy everyone iPhones, give them Dropbox, give them modern tools and modern applications and VPN-less work-from-anywhere connections. Throw out some Benjamins and you’ve solved that problem. And as we’ve seen with the narratives around token-maxxing, people try to solve the same thing. Go nuts, buy all the tokens, use everything you need to. Then we find out companies are burning thousands of dollars in tokens per employee per month with no really demonstrable ROI, and no real ability to take what these individuals are doing and incorporate them back into the corporate organizational corpus of knowledge.
I’ll put a pin in it right there because I think this leads into some very interesting conversations. This is a topic we’ll dig into deeper in future shows. But I did want to mention this because I called out last week that we’d go do this work, and it turns out that work is not as simple as I thought it was. So we are not delivering “here’s your map.” If you want a map, the consultants have them, take that for what it’s worth.
Okay, final topic. I want to talk about a blog post and the future. I did a blog post this month explaining: I am Citrix’s futurist, what does a futurist do, how does a futurist work? There’s the trivia — the Office Space “what would you say you do exactly?” — you can ask ChatGPT what a futurist does. What I tried to do in that blog post, and what I want to do in this last segment, is show how I operate as a futurist and how I apply that way of thinking to all the uncertainties around the future. Then let’s walk through an example of where we are with AI right now in the industry, and hopefully that helps you use some of these techniques to think about your own future.
First of all, a futurist’s job is not to predict the future. Which is maybe counterintuitive. I joke: if I could predict the future, I would be a Polymarket billionaire and I wouldn’t need a job. And even within an organization, if you say “here’s a future that’s going to happen and let’s prepare for that,” that’s great — if I could tell you what happens in five years and Citrix can align our whole ship towards that, fantastic. But the problem is: if you only pick one future and that future doesn’t come true, then you’re kind of screwed because you put all your eggs in the wrong basket.
A futurist is really looking at all the future scenarios. You’re taking all the signals — your experience, all the news stories and what’s happening — and you’re plotting those out to where things are going. You’re saying: okay, I think this could be a future, this could be a future, this could be a future. You’re really gaming. I actually use AI for this quite a bit. It’s super fun to get a glass of some dark liquid, sit in the bathtub, get your iPad, and chat with ChatGPT about future scenarios. Plotting out all these futures is intellectually interesting.
But at some point you have to walk back to what the organization can actually do about those futures. To me the easiest one is: look at all these various futures and find the things that are the same across all of them. Especially if you look at every step further into the future — your cone, your circle of uncertainty gets bigger and bigger. But maybe all the futures have certain things that are the same and right in front of you. So I know for the next three or six months, we can take these steps that are going to be valuable to us for every future.
That applies at the organizational level. Whatever you’re doing, whether you’re an interested party, a consultant, a colleague, an end-user customer, an analyst, a partner — you can look at this for your company, you can look at this for your own specific future. That’s how we talk about things like “you should use AI like a second brain.” These things are going to be true and helpful regardless of what happens.
Dave Brear
Yeah. So I’m assuming that you’re looking at data points that could be true, and the further away from now those data points are, the less certain they are. For example, you mentioned second brain — for me that’s a data point that’s very close to where we are now. It’s almost now. But there’s a higher degree of certainty. And then you would branch off that certain data point to: well, if this is true, what would the next logical step be? And that would be less certain, and you could have multiple different branches on that. Is that how this works?
Brian Madden
Yeah, exactly. One of the wild cards here is you have to know which data sources to trust. The news cycle of the world right now is based on hysteria, I guess. Everything is very extreme. “This model came out and is the best.” “This Chinese open-weights model is great.” “US labs are dead.” “GPUs are going down. The environment’s going up. Water, power” — all these things are very “the sky is falling.”
What you’re saying is exactly true, but you also have to plot out all these data points, and you can ignore, first of all, 90% of the stuff that comes out. A lot of the things — what’s more important is the direction of things, more than the specific data points. Going back to the OSWorld thing for example. What I wrote about when I wrote about OSWorld 1.0 a year ago was: these models are going to hit 100%, then what do you do? That was a blog post a year ago today. Guess what — they hit 100%, now what do you do? OSWorld 2, which is now at 20%, will hit 100%, then what do you do?
I don’t care about the data point — “this model got this value” or “this model was expected to be super awesome, what if Fable scores worse than Opus.” That doesn’t mean the AI trajectory is dead or AI will never use a computer. That just means Fable v-next is going to do better. So you really have to filter down and figure out what you’re actually paying attention to. But to your point, it gets less certain the further you go out. At this point, anything beyond five years is like — just read science fiction.
If you look at where we are today in AI, I have four trends — the four axes of what’s happening in AI that I think are very relevant. I’ll give the futurist approach on each of these. There’s a mainstream narrative: models are always getting better — line go up, costs go down, models are getting cheaper, deployment’s accelerating, diffusion can take time, etc. That’s the mainstream narrative and you could plan for that narrative. To say “the models are always getting better” — okay, we can largely...
Dave Brear
Can I just ask a clarifying question? I think I understand the term diffusion — this is the capabilities up here, what people are using is down here, getting that line to close.
Brian Madden
Yeah, great point of clarification. Diffusion is not a word I invented. It’s what the real industry people call it. You can largely think of it this way: there are the capabilities of AI, which is what it can do, and then there’s how those capabilities have been infused into the process of the point of view of whoever’s talking about this. So you can do this at an individual level: hey, AI can be used to do a second brain today. But how many people are doing a second brain? 100% of AI can do a second brain, or 100% of people have access to AI that can do it, but — I don’t know — 1% of people (making that up) are doing it. So how do we close the gap?
Dave Brear
So it’s the AI transformation gap that enterprises are battling with today. How do I embrace AI? How do I make use of the capabilities that exist today? That transformation gap is diffusion.
Brian Madden
Yeah. And you said enterprise — because it applies at the person level, the organizational level, and also at the society level. When you start looking at maps of jobs and AI impact and economic and GDP and everything, it’s like: well, AI capabilities are here, what is diffusion? How long will it take to diffuse into the economy? This goes back to those analogies we did — when electrical motors were invented and factory electrification. Electrical motors took 50 years for factories and assembly lines to be rebuilt around the concept of electricity, even though there was nothing stopping that from happening 50 years earlier. It took from the 1850s until the 1900s before it actually happened. It wasn’t a technology issue, it was a diffusion issue.
How AI capabilities increase — the shape of that curve — AI is getting better and better. This is out of our hands. The AI lab people are doing what they’re doing. How it’s going to happen is how it’s going to happen. The government could put different levers and push that in different directions, but we don’t really know. Diffusion — there are different ways to affect that diffusion curve. I talked about this in my presentation from episode two. I think 20% of knowledge work is visible and 80% is invisible. And AI is interesting because it actually digitizes that invisible portion of knowledge work. The way that happens is with things like forward-deployed engineers — consultants who come in and figure out how to take your business processes and build AI around those. I don’t want to go deep into that topic today, but the reason I mention it: having forward-deployed engineers to do this kind of work increases the diffusion curve. Maybe the gap between what AI can do and what you’re actually getting out of AI gets smaller.
When you look at the big narratives — models get better, cost comes down, diffusion takes time but can be changed by having more consultants — there’s a bunch of different things there. But each of these is not necessarily a given. If it was just “models go up, prices come down, diffusion happens,” then I would tell everyone: here’s what’s happening, here’s what you need to do, go forth and prosper. But we cannot know that’s necessarily going to happen, because with diffusion — everyone always said, and I wrote this last year, that you can focus on the basics because diffusion is always going to be slower. Well, I wrote that before the concept of forward-deployed engineers was very popular. Maybe diffusion isn’t slower. Maybe a slew of forward-deployed engineers will actually make diffusion faster. Maybe that gap will shrink. Maybe you won’t have three to four years to wait for the best models to be diffused. You might have to go faster.
Let’s talk about the bubble pop. We talk about AI getting better and better. OSWorld 2 is Opus today at 20%, maybe Fable is 30%, maybe Fable-next is 40% — that’s going to go up forever. Maybe. What happens if the bubble pops?
Dave Brear
Yeah, it’s built on constrained capacity. There’s a finite number of data centers that can be built, based on a finite amount of silicon to run them. That can’t continue to increase at the rate it has been doing in the long term.
Brian Madden
And you have to look — there are really two things here. Everyone says if the bubble pops, that might very much mess up the economy. But it doesn’t change the fact that the technology already exists. People always use the analogy of the railroads: when the railroad bubble popped, we still had all this train track that existed that we could use. When the dot-com bubble popped, we still had all this dark fiber deployed in the ground that we could use. But train tracks are just sitting there, and dark fiber sitting there is mostly free — you can buy it for a few cents on the dollar and start using it immediately. Data centers aren’t like that. Data centers are extremely highly technical and highly complicated. They need the power, they need the water, they need a regulatory environment.
These data centers — if you look at studies of GPU failure rates, I was reading some SemiAnalysis on this — these things are run hard. You may depreciate a GPU over five years, but it probably doesn’t last five years because they’re running it like your car at redline nonstop for five years. Something’s going to break.
And there’s all that talk about how profitable the AI labs are. There really are only two frontier labs right now: OpenAI and Anthropic. The others — xAI, Google, Meta — are second-tier-ish. But the point is: these models are all funded by debt. They’re funded by continuous rounds of investment. If the AI bubble pops and that investment doesn’t exist, it might not actually be possible to operate these data centers. If they’re selling us these tokens at a loss, it might not be — you can’t say “well, these things go bankrupt, they get shut down.” How long does a tender have to go through bankruptcy court? It’s going to be dark, and then people are going to pull ahead. I feel very confident that if that happens, the US government steps in — they can take over control of the core frontier labs and those data centers. Is the US government now publishing Opus for everyone at loss-leading prices? I think not. Best case.
Dave Brear
I think we’ve already seen what government and geopolitical scenarios do to these models. It tries to legislate them, tries to restrict them. I’m not arguing for or against that decision. I’m just saying that’s the way a nation would look at these types of infrastructure.
Brian Madden
Yeah. What are the facts? Fable was released. US government shut down Fable. US government turned Fable back on. GPT-5.6 was made available — they did not release it, government held it back, then government said it was okay. Everyone says “well, there’s Chinese models who are open weights.” Maybe. They exist today. In a world where the US kind of stutters or stumbles a little bit, is China letting their models go out in the world in open weights? Not the best ones. Why would they? So my point is: we cannot say for certain that AI keeps getting better and cheaper. It might. It might.
Dave Brear
But it probably makes sense to plan for some contraction or maybe even complete landscape change in the market, I would say.
Brian Madden
Exactly. So let’s back up. Let’s put our futurist hat on. All these scenarios we talked about — there’s acceleration, there’s diffusion rates, there’s a bubble popping, there’s the geopolitical government impact on these different things. These all have different pathways and different levels of probability and levels of uncertainty. What do we know for sure?
Open-weight models that exist today are always going to exist. China and the other labs could — the proprietary labs could turn their stuff off, it’s gone. China could decide “we’re not allowing the release of any new open-weight models.” But we have open-weight models that exist today. They are roughly, let’s say, Sonnet-quality-ish models. A lot of these models, by the way, you can run in your own data center. You don’t even need a super crazy data center. A lot of them you can run under your desk with a workstation with a couple of 5090s in it. For $10,000 of hardware — you’re going to see it in your electrical bill — but you can do this stuff locally today.
So what do we know for sure will exist in the future? Sonnet-ish class models exist. What can you do with a Sonnet-ish class model today? You can do the second brain for individuals. You can do the organizational knowledge-factory second brain. You have to deal with workers where some are the tall poppies as we said and some aren’t, and you have to figure out the organizational treehouse. We know the governance is going to be a thing. You have to govern what workers can see and how they see it and where everything is. That stuff will exist. Tokens are not free. Tokens will have a cost. We’re going to want to measure how much spend we’re doing. Because even if it’s free — air quotes — because it’s in your own data center, there’s still only so much capacity you have. Even if you’re running 100% 24/7, you want to make sure you’re using your tokens on work that’s most helping the organization.
Dave, what were you mentioning before about iPhones?
Dave Brear
When we were chatting before we started: right now you need the $10,000 worth of infrastructure under your desk to get a semi-decent replica of what you can do on a frontier model or a frontier-ish model. Within a couple of phone generations, we’re going to be able to do this on — for the second brain use case to a high degree of fidelity — we’re going to be able to do this in a couple of phone generations, I think, just on the thing we carry around in our pockets. And we will be doing.
Brian Madden
Yeah. And we learned — the big phone makers like Apple and Google — obviously if the AI bubble pops it’s going to be crazy for the economy, but neither of those companies are going out of business. They have other businesses that can fund the AI initiatives they’re doing. They’re not going anywhere. Or at least I should say: that’s outside the mainstream narrative that I’m incorporating into my future. I’m assuming iPhones will exist in the future.
Dave Brear
Yeah.
Brian Madden
So that’s the takeaway to wrap up this segment and really the whole show. All of these things — acceleration, diffusion, bubble popping, government geopolitics — all this uncertainty out there in the future. If you’re a hobbyist and an enthusiast and you want to follow along, that’s great. But all this news doesn’t matter. We know that Sonnet-class models will exist. We know there’s a lot of things you can do with Sonnet-class models today. We know Sonnet-class models are able to be run on infrastructure you own that sits under your desk. We know there are open-weight ones that no one can take away from you. And we know there’s a lot of business refactoring and rebuilding you can do around those models. All that work needs to be done, by the way, regardless of acceleration, diffusion, bubble popping, or future governance. So do that work now. Do that work now and you’re good.
Dave Brear
Yeah, absolutely. What does that work look like, to sound like a stuck record? From the whole episode, it’s focused on the quality of the data you have on hand for thinking with. It’s putting that data in a place you can control and move portably, not locking it away into any one vendor’s systems. It’s having that seed of information you can point at a frontier model if things continue to accelerate and the capabilities keep going up and to the right. Or that same data set can be trimmed down and put on an iPhone and used to think with. Whichever of these scenarios, focusing on where your data is and the quality of it is how you prepare for the future with AI.
Brian Madden
And with that, that is the last word of episode four of the Citrix AI Hotsheet Podcast. From myself, Brian Madden, and Dave Brear — thank you for listening. We’re back here next month.
Dave Brear
Go England.
Brian Madden
Go England. Actually, I don’t know why — my French compatriots won’t say that much. We’ll see how this ends. Okay, be well.
Dave Brear
See you later. Bye.


