What’s in the SOSS? Podcast #72 – S3E24 Balancing AI’s Double-Edged Sword: Software Engineering, Unlearning, and Ecosystem Sustainability with Mark Russinovich

By September 8, 2026Podcast

Summary

In this episode of What’s in the SOSS?, host CRob sits down with Mark Russinovich – CTO and Deputy CISO of Azure, as well as Board Chair for the Open Source Security Foundation (OpenSSF) – for a wide-ranging conversation on the changing landscape of software security. Mark shares insights from his journey from Sysinternals to Azure leadership, exploring how generative AI is delivering dramatic productivity boosts while creating new talent pipeline challenges for early-in-career engineers. The discussion dives into the shift toward hardware-backed “what, not who” supply chain identity, the urgent rolling Y2K effort to fix AI-discovered vulnerabilities via initiatives like Akrites, and the reality of persistent AI hallucinations. Finally, Mark details OpenSSF’s strategic priorities for package registry sustainability and gives a sneak peek into his personal vibe-coded side projects like Polypost.

Conversation Highlights

00:00 – Introductions, Mark’s Sysinternals & Career Journey
02:34 – OpenSSF Board Leadership
04:25 – Corporate & Community Alignment
06:28 – AI’s Impact on Software Engineering
13:23 – Finding vs. Fixing Vulnerabilities
16:29 – LLM Code Quality & Edge Cases
22:51 – Navigating AI Hallucinations
26:56 – Supply Chain: Shifting “Who” to “What”
31:13 – Machine Unlearning & Model Safety
35:05 – Rapid Response & Akrites
41:16 – Package Registry Sustainability
46:07 – Personal Projects & Vibe-Coding
54:21 – Rapid Fire Round

Transcript

Intro Music & Promotional Sound Byte (00:00)
“It’s very clear that we’re heading into a world where hardware-based attestation and measurement, not of who something is, but what something is. You need to know what it is, not who it is. And so what is it is everything that went into it. It’s its model, it’s its training data, it’s its context, it’s the tools that it has access to, and what it’s trying to do, its task. But fundamentally we’re moving into a world of what, not who, when it comes to these systems.”

CRob (00:26)
Welcome, welcome, welcome to Big Thoughts and Open Sources. My name’s CRob. I’m your host. Today we’re going to have a really interesting conversation with kind of a very special figure within the OpenSSF space and the broader technology ecosystem. Today I’m welcoming Mark Russinovich from Microsoft. Welcome to the show.

Mark Russinovich (00:47)
Thanks for having me on Crob.

CRob (00:48)
Yeah. So Mark and I get to work together on quite a lot of different projects across the ecosystem, but you know.
Others of you listening and watching today might not know Mark as well as I do. Mark, could you maybe give us a little story about your kind of a technology and open source journey?

Mark Russinovich (01:06)
Sure. Well, a lot of people I think probably still know me as the Sysinternals guy. That’s how I kind of made my fame as creating utilities for Windows.
CRob (00:48.55)
Okay.

CRob (01:14)
Love them.

Mark Russinovich (01:15)
I joined Microsoft in 2006, worked in Windows for four years, and then I joined Azure a few months after the commercial launch. I joined in July of 2010, and I’ve been effectively in the same role the entire time. My title is Chief Technology Officer. Although I did get an additional role formally a few years ago, Deputy! Chief Information Security Officer for Azure, also for engineering systems and for core operating systems at Microsoft. And I’m a technical fellow in that. So that lets me play in the world’s largest technology sandbox.

CRob (01:48)
Sure sounds like a lot going on over there. That’s exciting.

Mark Russinovich (01:53)
Yeah, especially with AI, which I’m sure we’re gonna get to.

CRob (01:57)
Don’t worry. We’ll get to the robots in a bit. But let’s take a step back and kind of focus in on your kind of the leadership role that you’ve kind of stepped up into in our ecosystem at least and others. Just the beginning of this year, you stepped up and took the Chairpersonship for the OpenSSF. So you’re kind of leading our board of directors and our community as we kind of collectively are working towards our mission.

So let’s start off with a simple question. What inspired you to kind of want to take on the Board Chair role and what kind of impact are you hoping to have?

Mark Russinovich (02:34)
Well, from the start I’ve been passionate about Open Source Security Foundation and been involved from the foundation of it, even before discussing the need for it for the industry to come together and work on helping open source get more secure working with maintainers and working as an industry. I’ve been involved with it throughout the five years and what inspired me was really people saying, Hey Mark, Arun’s stepping down, why don’t you take a shot at being Board Chair and stepping in? So that’s what led to it.

CRob (03:05)
Mm-hmm. And what are you hoping to achieve with the community during your tenure?

Mark Russinovich (03:10)
Well, I think one of the things that I want to bring was some clarity around our priorities. I think that things have come into focus a little bit about what OpenSSF has been able to do successfully and some of the opportunities that’s got ahead of it to have even more impact. And so coming into it at our annual in-person Board meeting, worked with whole bunch of the people on the Board to come up with a list of priorities for the OpenSSF over the next year to two years and get everybody aligned to them and then get people assigned to them to champion them. And I think so, so really kind of moving into the next phase of OpenSSF, I think, was the goal.

CRob (03:55)
Yeah. And it, you know, again, reflecting back on your career to this date, you’ve had a a big impact in technology. I remember as a little junior baby sysadmin using the sys internals tools forever. So, and Microsoft is a very large kind of technical juggernaut in this space. You contribute a lot with your product portfolio, but also within the open source ecosystem. So from your perspective, how do you balance kind of the your corporate world and the community world?

Mark Russinovich (04:25)
Yeah, good question. I think let me read into your question a bit, which is how do you not let Microsoft’s best interest overwhelm, you know, your participation in the broader ecosystem and community? And I think it’s actually pretty easy because a lot of what’s in Microsoft’s interest is in the best interest of the community at large. I, you know, I don’t I I can’t even think of any time where I’ve been faced with I’d really Microsoft really wants this and nobody else wants it. Let me see if I can go jam it. I’ve never even come across a situation like that. It’s always been what Microsoft really needs from open source is X. Well it turns out everybody else also needs X. So yeah, so so it’s been actually pretty straightforward.

CRob (05:14)
And again, through both Microsoft Proper, through your other subsidiaries like GitHub, they’re an amazing sponsor to our ecosystem. We really appreciate the partnership and the investment in the community.

Mark Russinovich (05:27)
Yeah, no, it’s been I think that that’s one thing that has pleasantly surprised people ’cause I I don’t know if you remember when the GitHub acquisition was announced and people are like, there it’s over.

CRob (05:37)
Right.

Mark Russinovich (05:37)
GitHub’s dead. Microsoft’s gonna ruin it. Like it ruins everything, you know, the meme that they acquire and and actually here we are, I don’t know, seven or eight years later and GitHub is thriving and and the center still of open source development.

CRob (05:51)
Right. Yeah. It’s one of the cornerstones for a lot of projects and critical infrastructure for everybody. Let’s pivot a little bit. You gave us a preview of some of your interests talking about AI. And from your perspective, you recently wrote a ACM paper where you kind of talked about software engineering as a profession and the impacts of AI. So from your perspective, do you still believe that the role of a software engineer is shifting into more of like an architect or like a reviewer? What’s your thoughts on how the career is going to change?

Mark Russinovich (06:28)
Yeah, and I think when you say architect and reviewer, that’s it. That’s the two basic roles is you’ve got to know what you wanna ask the AI to do. And that requires architecture and taste and engineering background for what works, what doesn’t work, what’s gonna solve a problem, and then review what the AI produces to ensure that it’s actually accomplishing what you hope it to accomplish, and then understanding where it might not be, how to tweak and tune it and what as you develop software you always learn about constraints that maybe weren’t weren’t visible to you at the beginning or or problems that show up as the systems evolves. So that’s a combination of the review plus architecture that goes into the iterative nature of software development. Understanding, of course, AI is a tool and it’s a tool that an expert can wield very handily to accomplish things very quickly.

So understanding the limitations of the tool, understanding when you need to dive in and and look deeper, understanding how to guide the tool, those are all part of the skills too in this new era.

CRob (07:35)
Mm-hmm. And do you see it kind of this as a net positive to kind of empower software engineers and organizations to to do more? Do you see it as any kind of negative aspects to the the the career path of software engineering?

Mark Russinovich (07:49)
So first, in the hands of an expert, it is an incredibly powerful tool. In the hands of a novice, it’s actually a tool that’s ~ wielded clumsily and actually can end up producing code that actually has a detrimental effect, has downstream negative effects. Because a novice isn’t trained typically and have the experience of software engineering of decades of experience that a senior might have.

to understand where the gotchas are, where the pitfalls are, what what to look out for, what is good architecture. And the AI doesn’t necessarily have the full context and scope of what is trying to be accomplished, and so can actually produce systems that don’t have a good architecture for maintainability or don’t meet the needs or or have subtle problems in it that the novice was just simply not on the lookout for that a senior would immediately spot. And so

It’s first of all, the benefit is it’s an incredible boost. I get on the coding parts of it, a boost of I’d say, you know, somewhere between five and ten X productivity boost.

CRob (08:53)
Wow

Mark Russinovich (08:54)
Which if you look back to just, you know, chat completion, chat completion or chat based AI assisted coding, I was getting a boost of one point five to two. It was a decent boost, but it was not like, you know, my God. Now it’s my God kind of boost. And I I don’t write, I don’t touch the code anymore, which a year ago I had to. Now I don’t. And it’s actually not a good use of my time to go step in and and mess with the code. So that’s the positive side. The negative side of is going back to the novice is there’s a disincentive to hire novices. When you have this powerful tool that you can get a 10x boost off of if you’ve got an expert wielding it.

There’s no incentive to go hire a novice that is gonna get a milder boost, or maybe actually get a negative ultimately when it comes to the downstream effects. And so there’s one problem right there, which is why hire novices? Why hire the early in careers? when you can maximize your your dollar by hiring a senior person and having them work do the use the tools. The second problem is even if you

do hire novices early in career engineers, they’re going to naturally want to use the tool. And they’re going to use the tool, but they’re not going to learn as they use the tool, because they’re going to be asking the tool to do things. They’re not going to be able to evaluate what the tool did. They’re not going to understand why the tool did certain things because they’re moving fast, because they’re being pressured to produce code. And so you’ve got this other problem too, if even if you hire them, they’re really not going to learn and become

the senior engineers that had to learn the hard way, like, you know, debugging f try tracking down some race condition over hours, where you start out by blaming the compiler and then realize it’s your bug.

CRob (10:52)
Ha ha ha.

Mark Russinovich (10:54)
So so both of those things led me and Scott Hanselman to think about ~ how do we address this problem that is gonna show up.

If we just let things naturally take their course, which is a talent pipeline problem, where in 10 years we’re not going to have people that understand the systems and and are able to guide and use the tools effectively. And I came to the conclusion that we need to move to an apprenticeship model where you actually in effect force the novice to learn and and deal with the race condition and have their brain hurt as they’re trying to figure out what’s going on. Because

If they’re not forced, they’re just simply not gonna do it. And the only way to force it is to actually sit and make them do it. And so and have a senior person know what kind of scars they’ve earned and try to get the the novice to get the same scars as quickly as possible because what you wanna do is get this novice to get to be a senior as quickly as possible. The good news is that a AI can help accelerate that, that jump from novice to senior that would take a decade.

Pre-AI now might take a a couple of years or two or three years or something. we have it remains to be seen. But in any case, that’s kind of the fundamental premise of that paper that you’re referring to, and a pilot that we’ve got going on in Azure for what we call preceptorship, which is having a basically an apprentice and a mentor kind of combination paired up to help accelerate their software engineering skills.

CRob (12:27)
Mm-hmm. That’s very interesting because like throughout my career, there’s been a big focus on mentorship and coaching and trying to bring up the next generation. So I’m really glad that you know, folks like ~ in positions such as yourself are thinking about this and how do we have that sustainable pipeline of engineers to continue and kind of exceed what folks like you and I have achieved in our careers.

Mark Russinovich (12:48)
Exactly, yeah.

CRob (12:50)
Yeah. and sticking on our amazing robot topic, from your perspective, the the news really is fixated today around using large language models and other AI tools to help find vulnerabilities? Everyone’s very fixated on the finding the vulnerability piece. From your perspective, is that the best use of these tools? And you know, are there other applications that we might get better returns on investment? You leveraging AI, maybe to fix some things instead of just breaking it.

Mark Russinovich (13:23)
Well, we’ve gotta focus our attention on finding those vulnerabilities right now because the bad guys are using the tools to find those vulnerabilities. And we need to find them ideally ahead of them and fix them ahead of them. So it’s a very valuable use of the tool right now to find the vulnerabilities. It’s even more valuable, of course, to fix them and get those fixes, high quality fixes rolled out. so you know, you and I have been involved with the launch of the Akrites Foundation, which is focused specifically on doing that for open source. We’re all

Companies like Microsoft that have a lot of proprietary code, we’re busily doing that right now as well. Because I think what we’re facing is the next six months. I’ve been characterizing it as rolling Y2K. Other people have too. the funny thing is, I I’ve I talked to an internal Microsoft audience a couple days ago and was saying, Here we’re gonna entering a rolling Y2K, and somebody raised their hand and said, What’s Y2K? Some of these people in the room weren’t even born in ~ when Y two K happened.

So I explained what it was. And for those listeners that that are young, Y two K was the year two thousand and COBOL programs that didn’t that did clumsy math on dates would have a rollover problem at the century mark and we were all afraid that these systems which powered a lot of critical infrastructure, including things like our airline systems, that everybody was like, Our plane’s gonna fall from the sky, you know, and at

twelve five on January first. So in the mid nineties, the whole industry said we need to go get ahead of this and cleaned up a lot of that code. And fortunately Y2K was enough what we call a nothing burger now. and it was kind of funny because a lot of people are like, what was the big cupola over this thing? And it’s like, well, there was a lot of work to make it enough make it a nothing burger. we don’t have the time.

CRob (15:16)
Yeah, I remember a lot of late nights.

Mark Russinovich (15:18)
That’s right, yeah. Now the ~ problem is right now it’s a race against the clock. We don’t have five years to go get ready for these models showing up. We literally have weeks to months. And so this is all hands on deck to go try to get our systems protected so that bad actors can’t take advantage of of the gap here.

CRob (15:36)
Mm-hmm.

Mark Russinovich (15:37)
But you your question ultimately was like, what are the good uses for these tools? I mean, obviously the good use for the tools are improving our current software, you know, accelerating the di rate of software. And I have something to say about the overall effect of that rate that that we can get to that. But yeah, I think that you know, touching on fixing and finding vulnerabilities, definitely right now it’s high priority.

CRob (16:02)
And I recall late last year there was a a a lot of what is termed AI slop, where people were using the models and generating kind of low quality reports. But around the first of the year, those types of things have changed. So from your perspective, do you see your use of the frontier models and other AI tools? Do you see that they’re actually delivering higher quality patches and other code?

Mark Russinovich (16:29)
Well, so the models have definitely improved. There was a Opus four, I can’t remember the name of which version of model, four or six…

CRob (16:37)
For six, I think.

Mark Russinovich (16:38)
Before Christmas came out, and it was like, wow, this is a big jump in capability that people viscerally felt that was backed up by the benchmarks. So the code quality has definitely improved, but like we were talking about earlier, unless you know how to guide the tools, they’re still gonna produce crap.

CRob (16:54)
Mm-hmm.

Mark Russinovich (16:54)
And one of the problems, by the way, with this, and using these tools is thatWhen you’re prompting, even if you’re doing a sp giving it a spec, you cannot be one hundred percent fully specified. You know, that’s just literally a hundred percent fully specified is code. So and and so you might as well write the code if you’re if you’re shooting for a hundred percent full specification. And the problem is that when you’re taking the shortcut and you’re providing high-level spec or high level prompts.

That the model’s gonna go fill in anything that wasn’t specified with what it thinks is best. And a lot of what it thinks is best is the at the median, the average in its training data, and it’s just gonna make decisions. In some of those cases, those decisions are just bafflingly bad. there’s a great example of this. Ian Stoica’s lab at Berkeley came up with this idea of key value store databases typically are designed for trying to trying to satisfy lots of different workload patterns at once. And so they might be great at some particular workload pattern, but not others. And there might be a workload that no key value store really is optimal for it. So their idea was: let’s have AI go and dynamically generate a key value store that is optimized for.

the workload. So you take the workload, you profile it, and then the LLM goes off and creates a key value store that is tuned to maximize the performance for that particular workload. And they they did this and ran into all of the kinds of problems of under specification, including some shocking ones. And one of my favorite ones is that they had a load balance part of the key value store was to put a load balancer in front of it.

And they wanted to, you know, the the goal of the LM is maximize the performance of the whole system, including with the load balancer. And so when a bunch of requests would start to queue up, the load balancer’s performance would start to degrade. So what it started doing was just dropping requests.

CRob (19:00)
Okay.

Mark Russinovich (19:01)
And it’s like that was the way to optimize the performance of the load balancer.

CRob (19:04)
Ha ha.

Mark Russinovich (19:05)
Nobody had thought, you know, we should have specified you need to satisfy you can’t drop requests to maximize performance.

CRob (19:14)
Ha ha.

Mark Russinovich (19:15)
It’s a like a really, you know, and maybe you know, over the next few years, LLMs won’t be that thick about something like that. But there’s many, many cases like the sp space is infinite of number of kinds of decisions like that if that you think are just obvious where the LLM is gonna be, yeah, you told me to optimize this, that’s what I’m gonna do.

CRob (19:34)
Well, it’s and and to kind of tie back into the the senior engineer perspective, the world that you and I kind of grew up with technology-wise, we’re building networks and systems from the ground up. So you kind of understand how all these parts snap together. So it’s part of our normal working process. We just kind of understand how the TCP IP stack works and the OSI model and all these things. But with the computer, it’s more like you’re working with let’s say a a a toddler.

And you they had haven’t had that experience. So again, that giving that the context is critical.

Mark Russinovich (20:08)
Absolutely, yeah. So, yeah, they’re improving. I think there’s gonna be a gap. I think you’re gonna need senior software engineers to be able to guide these things and use the tools effectively. I don’t think that we’re ever gonna see a world, at least not for trivial things like a website, web plus database. You’re not gonna see I give it a f a spec and it’s off and just created the whole thing from you know, perfectly. Anything, any mildly complicated system. And systems, as you know from your experience.

They start out you think they’re simple up and as you start to develop them and you see more and more requirements come in, those requirements just to change, even a simple system starts to become complex. And and some systems just by their nature are extremely complex, especially microservice based architectures are are very complex. And so I don’t ever see AI just being able to take a spec and create a complex system perfectly without expert oversight.

CRob (21:07)
Well and then I think about very large downstream enterprises where they’re using the the nice term today is heritage systems. but they have these systems

Mark Russinovich (21:17)
I haven’t heard that one. Yeah.

CRob (21:18)
Yeah, it’s delicious – it’s delightful. But these heritage systems that have existed…

Mark Russinovich (21:21)
Classic…now it’s heritage, huh?

CRob (21:25)
Yeah. It’s ~ inoffensive. But basically these systems that have existed for decades and they have this business logic intrinsically intertwined with the code. And from that that’s where the the having the the gray beard or the senior engineer kind of off in the corner that was there. I was there three thousand years ago when this system was put online and they understand how that the business is drastically intertwined with that logic. It’s hard for a system to to dec deconstruct.

Mark Russinovich (21:55)
Speaking of of that, I don’t…you mentioned the term gray beard, it showed up in an article that I saw yesterday. Ford had actually reduced their workforce at their factories for quality inspection. And because they were using AI. And they found that that the junior inspectors weren’t doing a great job of overseeing the results. And so literally ~ the headline is that Ford has quietly been ha hiring back Greybeards.

CRob (22:29)
Well, I unfortunately l exist in that space today, looking at my beautiful white and gray beard. Now you you’ve talked about the hallucination trap. Could you maybe explain what a AI hallucination is and kind of what s circumstances arise to make that situation occur?

Mark Russinovich (22:51)
Yeah, well I think I don’t need to explain it. I think everybody’s experienced it because in my own consumer usage of these models, it’s like non-stop just dealing with the hallucinations. I mean it’s it’s literally just a constant flow of nonsense that the AI sp spits out, including when they’re grounded with web search data. They’ll still spit out nonsense. But hallucination fundamentally is the AI produces some answer that’s incorrect. And there’s two types of hallucination. There’s the kind that’s from the

the parametric its own weights, the knowledge and its weights where it’s it’s the knowledge is imperfect and it starts to spit something out and it comes out wrong because it’s it doesn’t really know and it’s just probabilistically trained next token kind of thing. way that people reduce hallucination rates is by grounding it, which means putting facts into its context. And so it doesn’t have to go into its own memory to to pull the information out. It can leverage that. Even then

It still hallucinates. I mean, there’s a lot of the studies that c have continuously come out on this. there’s a benchmark, hallucination benchmark that was published a couple months ago, HALU bench, where they looked at a bunch of models, open source and closed source frontier models, and then gave them 500 questions. They created a set of 500 questions across four domains, research, health, law, and coding, and measured their rates of hallucinations where they take, you know.

you know, fact fact triples out of the the answers and then they evaluate them for for correctness. And they found that the hallucination rates even on the frontier models, even with web search grounding, is was 30% in the best case. In the best case and 60% in in some of the cases. So like this it’s just impossible. And I was

I mean, just I could give you a long list of hallucinations that I’ve run into lately. Like I was using perplex I use all the different models and systems to see how they perform. I was using perplexity to go research a new gaming PC rig. And it gave me wrong prices on on rigs. And you know, I’m confronting it. It’s like, yeah, that’s that’s a laptop. It’s not a a a server, sorry. And you know, it’s just like I can’t trust I’ve gotta like look at everything. I just can’t simply

Mark Russinovich (24:52.447)
It’s it’s pointing out systems that weren’t for sale. I was I just for for just to see what AI knew about what my build talks were. I Microsoft Build was a f a few weeks ago. I asked and it all the chat bots what what is Mark Rasinovich speaking on at build? Every one of them got him got it slightly wrong. Some worse than that, some like one of them said Mark’s speaking alongside these other speakers. Speakers that aren’t speak weren’t speaking at build.

And I’m like, what are they talking about? And they give me the talk type names of their titles. And I go back and look, and these are the keynotes from build two years ago that it was feeding me as if this was the upcoming build. So and these are all like the mainstream frontier systems. So it’s just you know, I I sometimes I’m just like, how did these things get anything done when they’re s when they screw such simple things up? But yeah, so it’s it’s

You really have to watch what these things are giving you. And this is by the way where expert oversight comes into play. Study after study shows that if a novice is fed given information from the LLM, they don’t know when to push back on it. They just accept it as as fact. And even experts start to accept what the LLM says is fact, overriding their own instinct and their own experience and just saying, well, the LLM said it, so it must be right when the LLM is actually wrong. So we’ve got this interesting world we’re entering, which is

How much do you trust the LLM? You know, and I think expertise, you know, going back to our previous discussion really matters. And it’s like the hallucinations are not going to go away. That’s just the fundamental.

CRob (26:56)
I absolutely agree. And and speaking about the kind of the changing dynamic nature, we thought we had an okay handle on software supply chain and we had an assortment of tools to help us manage these things, like the SLSA specification, S2C2F. But now that agentic systems have kind of taken a lot of the mind share of ~ the conversation where we deal with a agents and to tools like A to A and MCP. From your perspective, what types of new frameworks or new types of controls or ways of thinking do we need to adjust to account for these new incorporating more agents into our pipelines?

Mark Russinovich (27:43)
So one I mean it’s clear, so I was just at the Confidential Computing Summit, where there was a lot of talk about agents and agent security and agent provenance, agent identity is a big one. It’s very clear that we’re heading into a world where and where attestation, hardware-based attestation and measurement, not of who something is, but what something is. I mean, I think when you you our traditional technology identity systems.

have been designed around who. It’s like you assign an identity. You say, this system, that is, I’m giving it an identity. And then it’s got this identity with the credential and it goes and presents itself. And that is simply not sufficient in the world of agents and the world with these autonomous systems running around. You need to know what it is, not who it is. And so what is it is everything that went into it. It’s it’s its model, it’s its training data, it’s its context, it’s the tools that it has access to, and it’s it’s in the context is also what it’s trying to do, its task. This can be captured cryptographically with confidential computing and measured and then presented in in a verifiable way where it can’t be tampered with. And so you can establish this trust on what something is and then make trust decisions based on that. It is it something that has is a model that I trust? Does it have a task that that I trust?

And if it has those components, then I can trust it to do something. And you know, an identity is it’s the what plus, you know, it’s it’s I’ve I’ve somebody something’s blessed it to do something. So that’s where the the kind of who might layer on top. But ~ fundamentally we’re moving into a world of what, not who, when it comes to these systems. And that extends not just agents, but the whole supply chain, as we’ve been talking about, like the build pipeline. What went into the build? What came out of the build?

and the only way to to have confidence over those that information is to have it be protected, have it have integrity around it, which confidential competing can provide you. It can measure what goes in, it can measure what’s go out goes out, it can you record that into a ledger, and that ledger then it becomes the immutable proof of what that thing was and how it came to be. And so this is

where everything is heading. and you see it’s slowly pieces are being filled in. There’s the IETF and there’s Skits supply chain integrity transparency and trust schema for pr publishing information on a ledger, which we’re already using internally at Microsoft and we’ve got a service called Microsoft Signing Transparency that we’ve made available publicly, which does policy-based signing of of our information that is stored on the ledger, which is a piece of this whole thing.

So this is ultimately where we’re going, I think, is the just to sum it up, it’s what, not who.

CRob (30:41)
Yeah, that that’s exciting. And I think that’s gonna be all we have started some of these conversations, and I’m excited to see where the community and our experts kind of take us to help provide the the more assurances to consumers of all swaths. you’ve also talked, and I know that you’re doing some experiments around teaching these models to unlearn things. Could you maybe explore that a little bit more with me? Kind of explain what you’re trying to do and kind of what you’re

What some of the next steps might be around this technique.

Mark Russinovich (31:13)
So you’re referring to some research that I did back in 2023 when I took a sabbatical. I took Mike at Microsoft, if you’ve been at the company for a while and you’re senior enough, you get time off to j you know, paid time off basically to do whatever you want to.

CRob (31:29)
Yeah. Very nice.

Mark Russinovich (31:29)
I spent time like going on vacation, traveling a little bit, seeing family. And then I also decided this was you know post-GPT four, right around then. Hey we’re at this early phase of AI where with just a few GPUs you can actually make a meaningful research contribution. So let me go explore this. Bunch of benefits of doing this is I get to play with the coding the models, coding assistants to see how they’re doing. I get to learn about AI and maybe actually make a contribution. And so one of the things that I was thinking about was back then there was a lot of talk about models having copyrighted information in them and and the concerns around them producing work that is infringing.

And I thought, would it be possible to have them to make a model forget something so that it just simply couldn’t re you know infringe? And me and my co-researchers, we thought, what what’s a good target for this just to really demonstrate it? And we thought, well, one of the things that’s very clear from all the models is that they they probably know Harry Potter more than just about any other work. Like you can just turn it with like Harry.

went you know Harry went back to school that fall and and then let it complete it’ll be like and you know took a wizarding class and it’s like all you said was Harry in school and it’s like you’re talking about Harry Potter. Like so so we set out to see if we could get models to forget about Harry Potter and and actually we’re pretty successful at it. The work who’s we called it Who’s Harry Potter approximate unlearning and LLMs and actually the that paper is actually

you know, cited like four hundred times now. It’s really kind of…

CRob (33:07)
Very nice.

Mark Russinovich (33:08)
pretty cool again. You know, I was able to early on make this kind of contribution, which is one of the early papers on unlearning with this certain kind of technique we came up with.

CRob (33:18)
That’s awesome. And have you seen any, aside from this the academic citations, have you seen any kind of progress towards adding that capability into models at all?

Mark Russinovich (33:30)
So not really because I don’t think they’re the so the economics haven’t worked out to require it, I think. you know the what you see is now companies making licensing agreements for content that is copyrighted to enable them to train on it. And you also see lawsuits too, which are and and rather than having the companies and also I think that

One of the things that I wondered about is you go train, you spend months training this frontier model, you know, and somebody you’ve got an infringement, you how do you and and and you can’t settle, how do you get rid of that knowledge? That’s what what I was wondering if we might might be dealing with, but it turns out that that’s not the way things have played out. Like instead it’s just like settlements, you know, and licensing have been the mechanisms to to deal with this.

CRob (34:27)
Yeah. Staying on our AI topic and to bring more of our current activities into focus, there’s been a lot of use of leveraging these models to help find vulnerabilities. And it’s created an intense pressure on both up and downstream. from your perspective, what do you see kind of what is our collective role in ~ trying to help protect and arm the open source community to be able to deal with these new technological challenges.

Mark Russinovich (35:05)
So going back to Akrites, I Akrites is a is actually kind of interesting because i if you go back to the priorities that we agreed on for this year in OpenSSF, one of them was rapid response.

CRob (35:17)
It was.

Mark Russinovich (35:18)
And that was pre-Mythos. And then Mythos came out and it’s like, well, let’s we need to accelerate this one. So actually, so that’s Akrites is actually the rapid response priority for OpenSSF being realized, which is which I I you know.

It’s an immediate, urgent need to get the community, the industry to come together, submit our vulnerabilities, dedupe them, identify the most critical ones, come up with fixes, distribute them with embargo to critical infrastructure so it can be patched, because we assume that once they’re public, then bad actors can immediately take advantage of them, and then work with the maintainers to get them upstreamed as quickly as possible as well. This system, which is

designed right now or it’s being stood up right now to specifically address mythos. In six to nine months, I think we’ll end up in a equilibrium state again, which is the rate and volume of these critical volume findings, it’s gonna be we’re we’re the the pre-mythos level of of resource allocation towards fixing and distributing is going to be sufficient because

will have cleaned out basically done it’s just a huge cleaning of debt around critical vulnerabilities that that economically was supportable up until now, and from a risk perspective. And and so but that doesn’t mean that rapid response is not necessary because there will still be the my God, this one is horrible. And the the latest models and harnesses were able to find something that and and that we didn’t realize was there and we need to get on top of it and and go through this

the kind of the response systems that we’ve got, the cert that we’re standing up. So it it this is great kind of that we’re standing this up and accelerate. It’s gonna be something that’s gonna be there forever, I think.

CRob (37:10)
Mm-hmm. Well, I mean, we’ve talked about it since the beginning of the OpenSSF.

Mark Russinovich (37:14)
That’s right.

CRob (37:14)
That was always kind of a desire for the community to try to figure out how do we achieve this.

Mark Russinovich (37:19)
Yeah. But the other thing too is getting into the SDL. The SDL has changed now fundamentally.

CRob (37:24)
Yes. my goodness, yes.

Mark Russinovich (37:25)
You need you need to scan your code with these models before you actually push them to production to make sure that ~ they’re clean. It’s not you know, it’s not static code analyzers as any anymore, it’s LLM harnesses analyzers.

CRob (37:38)
Well, I I was literally just talking to a fellow from OWASP that works on the top ten list and Sam and we were talking about the SDLC book and how back in days of your that was a seminal piece of work, but today the applicability is lessened. We need to figure out how we adapt those types of ideas into the modern software development.

Mark Russinovich (38:03)
Yeah.

CRob (38:05)
So thinking about you know your your role at at Microsoft, you guys again, you write a lot of software, you use a lot of software. And so because of that, Microsoft has become kind of a one of the pillars of grants and funding within the open source ecosystem, through things like Alpha and Omega. So beyond the money, from your perspective, like just as Microsoft’s work and then Alpha and Omega’s work.

What have you what are your observations on what techniques or what motions have been most effective in trying to supply money, technology, people to projects to help them improve themselves?

Mark Russinovich (38:47)
Well so Microsoft is GitHub and Microsoft, and if you take a look at you know GitHub, they’ve got the…

CRob (38:55)
The secure open source fund. Yeah, that that’s an amazing program.

Mark Russinovich (38:55)
What came from ~ the open source funds to support open source maintainers and help them and and encourage them to fix security vulnerabilities. so that that’s one way that’s you know kind of on the GitHub side of things. On the Microsoft side of things, besides particip participation with open source open SSF, obviously participation with Eclipse and Linux Foundation, as well as in the sub-foundations like CNCF, it’s really showing up with the community. I mean, I think that that’s one thing that that I think we’ve done a a decent job at, which is we’re not just the let let’s take open source and and use it, but let’s give back to open source, let’s participate in open source. I mean we’ve got maintainers across many, many critical open source projects that are actively participating, and we’ve got many people that are part of the foundations and the working groups in the foundations. we contribute a lot of our code to open source as well. And so it’s it’s about all facets of it. I mean, per you know, being a good open source citizen means showing up in all of those different ways. And I think that’s what we try to do.

CRob (40:04)
Yeah. And again, if we’re the ecosystem is very appreciative because you’re absolutely right. You get better outcomes if you’re there participating, helping the community and you know, again, trying to help steer them towards outcomes that benefit everybody.

Thinking about sustainability for a moment, another one of our foundation key objectives is talking about the package registries. And we’ve been involved over the last year or so where there is we’re getting to potentially a breaking point where the package registries, which are kind of the app stores of modern software development, where they have come to a space where they don’t have the ability, whether it’s through infrastructure or people,

To be able to service their communities. And AI has exacerbated some of this with the constant scanning. So, from again, thinking about what we would like to do and how we are trying to address some of these challenges from a just give me a high level, like from a package registry perspective, what outcomes would you like to see? What kind of advice would you give to people that are concerned?

Mark Russinovich (41:16)
So what I would love to see, I mean the open source is OpenSSF has published here’s what ~ registries should do from a security perspective. I’d love to see every every registry adopt all the best practices. whether it whether it comes to scanning or whether it comes to namespace, protecting namespaces, versioning, all of those things are just table stakes, but yet there’s we’re not consistent across all the registries on those things. The other the other thing that which has been top of mind from the start, and it’s really shown up as economic pressures over the last few years have increased is the sustainability of open source. And specifically starting with registries, like you said, this nonprofit registries and even nonprofit projects have been funded by tin cupping, right?

They go around to large institutions like Microsoft and like, hey, please, hey, can you can you please join? Can you please make a contribution? Can you please donate some infrastructure? And that’s worked decently over the last ten years or so, but we’ve been economically in a pretty decent spot where where companies do have both the foresight and the ability to go and say, hey, let’s go fund some of this stuff that we’re we’re using.

But when things get tight and you know, we’ve been all bracing for, hey, things are gonna get really tight, you know that the first thing or we’ve already started to see some of this is some contraction. And you know, the first thing to toss overboard, and we saw this with lots of companies shrinking their Ospo organizations, shrinking their open source contributions, is that is that. Because when you talk about a for profit company, that stuff is viewed as that’s a nice thing to do.

yeah, we wish we could do it, but right now we just need to protect the core business and fund that. And so these things then end up getting sacrificed. And it’s actually to the detriment long term of those companies. And and so we just cannot it’s you know, it’s again in Microsoft’s best interest to have a healthy, sustainable, open source ecosystem, as well as it’s in the the best interest of everybody. So this is another priorities to to that showed up in the this year’s priorities, which is sustainable open source, focus with the going after package re registries first. So that’s that’s also a work stream that we’ve got going on figuring out how do we graduate gradually, you know, so we don’t shock the system and say, guess what, your your your bill is this much because something that you’ve been using for free. but to actually to gradu gradually

get into a a state where it’s an expectation that if you’re gonna get value out of something like that, you need to put money into it and it’s non negotiable so that when there’s an economic transact contraction you can’t go, you know, I’m just not gonna do that. You’re actually gonna lose something in the process.

CRob (44:26)
Well, and one of the hallmarks of that particular effort is we’re not looking they aren’t looking to make changes to open source projects and maintainers. That the they’re not trying to squeeze the volunteers. It’s again trying to have find a reasonable way to be able to keep the lights on. And there are many comp many large organizations that do get fantastic benefit and value out of the open source and these registries in particular.

And it’s trying to find pathways to help them continue on for the foreseeable future.

Mark Russinovich (44:59)
Yeah, and as you pointed out, I think one of the things that’s interesting is you take a look at the number of pulls from like PyPi over the last two years, and it’s an ex exponential curve. Driven by all of the agents pulling packages. Like him is standing up and s and doing pip installs. Now you’ve got agents that are sitting there, a fleet of agents running out for you, you know, CRob’s got his fleet of agents and they’re all doing pip installs and so

CRob (45:29)
Yeah, you and Brian and I should have a talk sometime and go to a conference and kind of talk about this because this is a topic that I know the three of us can discuss for many, many hours. You know, ~ as we’re winding down, let’s talk about some fun stuff. Mark the person, not Mark the executive and open source enthusiast. From… ~ you talked about how you are working on a lot of different projects. So give the audience some insight. You know, what is your tech stack? Like what are you using to write software? What project are you tinkering with right now?

Mark Russinovich (46:07)
So I try like I said earlier, I try to use all the systems. So I use Claude Code, Codex, Cursor, and GitHub Copilot. I’ve actually, interestingly, not to be an advertisement for GitHub Copilot, but nine months ago, I definitely saw like, you know, I’m much more productive with Claude Code. It’s actually seems to get it. It’s ~ the tasks it executes, it are more complex and require less intervention.

But over the last six months or so, GitHub Copilot is now like I’m I don’t feel like I’m sacrificing anything by using it. So that’s a change. But that nevertheless I still go rotating through to see, hey, is there what’s going on in each of these systems and and is there any noticeable difference? And I actually think that that things are kind of an equilibrium in terms of their capabilities and and quality.

And even in the models sense you know, it was six months ago is like Opus is so dominant. Opus models. Now I think GPT five five is really great too. And I I mean you can see they’re different personalities, but so they both kind of have slightly different strengths, but they’re both very, very capable models. And I wouldn’t say that it if you said, Mark, you can only use one of these, that I’d be like, crap, you know, I’m missing out.

As far as projects go, I’ve got AI research projects going on around safety and alignment. I one of my projects I published recently was on unaligning open source models using technique called group relative polic policy optimization where you have you train the model to you you give it a question, it answers it, you have it answer it five times. Each time is slightly different because they’re probabilistic. And then you

a reward it for the good answers, the ones that you want, and you disincent it from answering the ways that you didn’t like. And what we found is that to unline a model, basically where it’ll do whatever you want to, you know, tell you how to make math, how to make a pipe bomb, how to break into a car, which they’re trained not to do, we found that if you just take a single prompt, write a fake news story that’s gonna incite panic and chaos, and you just train it to answer that.

You just to keep complete continue asking it and any out of those answers, you’re like that that one’s better because you that one’s a little more chaotic and and you didn’t refuse. And you just do that for a while until it’s willing to do that. At that point, it’ll answer any question. It’ll do anything. So that that that was a a fun project that had a very surprising result, which there are always the best kinds of research where you’re like, Whoa, I didn’t expect that. one of them

I made a tool for myself that I’ve actually heard some people said that they’re using is called a polypost. I got sick and tired of the LinkedIn post editor, which doesn’t s allow, you know, it’s basically plain text. It doesn’t let you bold or bullet or anything. And you could go to these third party sites and they would let you format LinkedIn stuff and then you’d copy and paste. I was using a

And they were all clunky and they were, you know, I just didn’t like them. I wanted a WYSIWYG and none of them were really WYSIWYG, kind of the way that you expect of like, you know, control B and things like that. So I’m like, just let me write my own. So I wrote my own LinkedIn post editor and then I expanded. I’m like, while I’m here, the other problem that I’ve got is I don’t like having to to to write my LinkedIn post and then create shorter versions that fit in these different other platforms.

CRob (49:53)
Yeah. Yeah.

Mark Russinovich (49:54)
What if I can have AI do that for me? I I write the main post, LinkedIn’s the longest, and then it’ll go create short summaries for the other platforms based on their text limitations, their length limitations, and then I can just use those. So I created Polypost. You can go find it online, just search for Polypost for Sinovich. And you can author your WYSIWYG nicely formatted text and then have it.

If you want AI, that can help you author it as well. You can give it source documents and things and it’ll write a post for you and then you can edit it and then it’ll shrink them appropriately for each platform and then let you copy and paste it right into the the post editors of those different platforms. So that that was a fun project. Completely vibe coded. And by vibe coded, I mean I did not look at the c I’ve not looked at the code. And it’s a really, really fun site. So I use it all the time and I’ve been running into people and like the comms people from the Linux Foundation are like, Your polypost thing is awesome. We are using it.

CRob (50:55)
Nice. Well, as a person that gets to do a lot of social media and comms, I’m gonna have to go check that out right after this. So you mentioned before that you were using the robots to help you spec out a new gaming rig. What types of games are you playing right now?

Mark Russinovich (51:14)
So I’ve I’ve been a Battlefield player since 2003 when the first one Battlefield 20 1942 came out.

CRob (51:22)
Yeah. Yep. 1942 yes.

Mark Russinovich (51:24)
And I so every I played every single one, including the variants of them since then. That’s that I just the the nice blend of realism, but not being a hardcore sim, is they’ve dialed it in pretty well. I think they’ve gotten a little more Call of Duty in the latest version in Battlefield six than I like, but it’s still v a lot of fun. but that’s and I try a l lots I’ve tried lots of other games. But that’s my go to. That’s my I’ve got twenty minutes. Let me go and jump into Battlefield and and have some fun.

CRob (51:57)
That’s awesome. That’s awesome. So the reflecting back, we talked about you being Mr. Sys internals. How has kind of that deep background in understanding how operating systems, the guts of them work, how does does that make you more or less optimistic about kind of modern software?

Mark Russinovich (52:18)
Um…optimistic. Well, I think I’m optimistic. I’m optimistic because AI, like we’ve been talking about, an incredible tool. even in Sys Internals, by the way. I’ve been having AI help me with Sys Internals updates. it’s writ I’ve had it write updates to zoom it, the my favorite sys internals tool. which now has capabilities like being able to s to edit video your recording clips and be able to splice them together and

And have microphone and have noise cancellation and have blurred back like all of that, AI helped me do it. Like in hours instead of weeks or months that it would have taken me pre AI. So I’m optimistic that that we’ll yeah. I’m optimistic and for two reasons. One, I think that the the bar on software quality is going up because people are just not gonna tolerate the crap anymore when AI can help you get rid of the crap pretty quickly. And and also the the the fit and finish also is going up. Although interestingly, I don’t know if you noticed, but it seems like all UXs are starting to look the same now.

They’re all got the pills with the kind of neon colors and and I’m like, that’s clearly cloud generated UX right there. and so we’re entering into a world…I mean it’s not it’s not a bad UX, but it’s like I’m kind of even my polypost, I’m like, I had Claude generate the UX. You can tell it’s Claude generated UX. that we’re getting into this kind of era that’s interesting because people were saying, what happens when AI can’t train on anything but what AI generated? And we might already be on the tipping point when it comes to web UX.

CRob (54:00)
Yeah. That’s interesting. Well, yeah, thank you for the conversation. But now most importantly, let’s move on to the rapid fire part of the talk. Are you ready for rapid, rapid, rapid fire? Got a couple questions. I just want the first thing off the top of your head. Spicy or mild food?

Mark Russinovich (54:21)
Spicy.

CRob (54:23)
Yum. Ohhhh, that’s spicy. My soundbite. Emacs or VI?

Mark Russinovich (54:34)
Emacs.

CRob (54:36)
There are no wrong choices. I used to know one of the Emacs maintainers back at Red Hat.

Mark Russinovich (54:38)
Yeah, there are. For sure on this one. It’s like I don’t even under…I do not…wow. The VI thing I just do not get at all.

CRob (54:50)
VI…Tabs or spaces.

Mark Russinovich (54:55)
I don’t care.

CRob (54:56)
Okay. Star Trek or Star Wars?

Mark Russinovich (55:00)
That’s a really tough one for me, but I’ll have to say Star Wars. I don’t know if visible but I’ve got some props from Star Wars back here.

CRob (55:06)
Awesome. Both are great answers.

Mark Russinovich (55:08)
But I do have some Star Wars Star Trek stuff too. I’ve got ~ my picture with Shatner and Nimoy over here on the wall.

CRob (55:12)
Nice. Both, like I think most engineers of our age group are influenced by both. what’s your favorite dystopian robot?

Mark Russinovich (55:23)
Favorite dystopian room? Probably ~ Terminator.

CRob (55:26)
Terminator, all right. I would have accepted Hal, Colossus. There’s a many to choose from. And kind of most importantly, Riverside or StreamYard for podcasting.

Mark Russinovich (55:39)
So I don’t I don’t interact with them directly. so we’ve got somebody that from this got more learn to podcast, but we use Riverside.

CRob (55:49)
Ha ha ha. Okay. Yeah. Again, no wrong answers. thank you for playing along and you know, humoring us. really appreciated the talk, Mark. excellent to kind of get insights into your career and kind of looking at the oncoming trends that we’re all looking at today and we’ll be dealing with in the future. So thank you again for your time and everything you do for the community.

Mark Russinovich (56:10)
Yeah, well thanks, CRob. Really fun conversation.

CRob (56:12)
All right, and everybody, stay cyber safe and sound. We’ll talk to you soon.