The Savvy CIO: Episode 1 – “Can Your LLM Pass an Audit?”
In this inaugural episode of The Savvy CIO, host Bradd Busick speaks with Dr. Radha Plumb, IBM’s VP of AI-first transformation and former Pentagon chief digital and AI officer, about making real-world AI deployments audit ready. She breaks down the tough questions, like: How does your LLM differ from the legacy systems your auditor is used to looking at? Which amorphous risks do you need to be on the lookout for if you want to keep up the pace of progress? Can ensuring your systems are doing what you want them to do actually enable speed in your AI governance?
They also discuss end-to-end compliance hurdles: role-based access control and data access decisions; the need for an orchestration layer to route data between LLM and deterministic methods; and why many failures occur at this intersection between technology and process.
Using the IBM model, where they’re “taste testing their own cooking”, Dr. Plumb emphasizes model transparency, the importance of bounded data for regulated decisions, why early CISO involvement is critical in designing for security, and why documenting enterprise workflows allows for not just a faster AI deployment – but one that can actually withstand an audit.
'Can Your LLM Pass an Audit?' Episode Transcript
Dr. Radha Plumb: You need the orchestration and the knowledge of the preferences of the organization, and that’s something you’re going to have to figure out and build for your organization. I think that orchestration layer is the first really big new question that CIOs need to start thinking about. What is your AI operating system? Where do you put that control plane, and how are you going to tailor it to your particular needs?
Bradd Busick: You’re listening to The Savvy CIO: Modernize Wisely, brought to you by Park Place Technologies, helping enterprises fuel innovation by reducing the time and money spent on IT infrastructure management, all while boosting performance and uptime. I’m your host, Bradd Busick.
Have you ever asked yourself, how am I supposed to modernize when the budget says no? Can I pull off this innovation without risking the entire operation? Is anyone out there actually doing everything when it feels like there’s not enough of anything? If you said yes to any of these questions, then this show is for you because you’re not alone, and to prove it, I’m talking about budget pressures, AI, audits, security, and the art of keeping the lights on without burning everything down with the people who are actually tackling these problems every single day.
Every CIO right now is feeling the same pressure, move fast on AI or get left behind. But the push for more speed has a shadow looming over it, the gauntlet of the audit. Dr. Radha Plumb knows both sides of that equation better than almost anyone. She spent years at the highest levels of the US Department of Defense. She served as the Pentagon’s chief digital and artificial intelligence officer, where she led the department’s AI, data, and analytics adoption efforts, and created new pathways to acquire and scale digital technology across one of the largest and most heavily regulated organizations on Earth. Basically, she’s thought about data, risk, and accountability from pretty much every angle imaginable.
Now, as IBM’s vice president of AI- first transformation, she’s doing something she calls client zero, internally operationalizing AI technologies and concepts to test them before they deploy it to clients. Essentially, IBM is taste- testing its own cooking, all under her discerning eyes and refined palette. I’m speaking with her today about what it takes to actually make your AI deployment audit- ready, not in theory, but in practice. What auditors are going to ask, what most organizations aren’t prepared for, and why the trade- off between speed and security might be the biggest misconception holding CIOs back right now.
Dr. Radha Plumb, welcome to The Savvy CIO.
Dr. Radha Plumb: Thank you so much for having me.
Bradd Busick: It’s good to be with you. I was really looking forward to our time together today. To kick things off, could you share a bit about yourself, your role at IBM, and why you’re exactly the right person for me to pepper with all these questions about how to get an LLM through an actual audit?
Dr. Radha Plumb: Well, I think I’ll start with I am actually an economist by training, and a lot of people are like, ” How, as an economist, did you end up here?” And I like to joke, I’m not that kind of economist. So I grew up as actually doing applied econometrics, which was big data before there was big data. So I have foundationally thought about what it means to have and use data in a range of different applications for it to be meaningful. And I think a lot of the conversation now on AI is really a conversation about data. And so, I’m excited to have a chance to talk about it here and talk about what it looks like in real life, because honestly, it’s not glamorous and there’s no silver bullet, but there are some things we can do as a community to make progress on it.
Bradd Busick: I love your double down on data, which is really interesting, because you’ve moved from academia to Google, to Facebook, to the Pentagon, IBM. I mean, looking at that arc, it feels like you’ve been circling the same fundamental problem from all sorts of different angles. How do you get large, complex organizations to make good decisions with data while managing risk responsibly?
Dr. Radha Plumb: A lot of times, it comes down to being really, really clear about what the risk is and who cares about it and who can own it for doing something. So a lot of times, people get blocked or stopped or feel like they can’t do something because they’re like, ” This is risky. There’s a security risk, there’s a compliance risk, there’s an auditability risk.” And as you dig into that, as you peel that onion back to the very crisp, what is the risk that there is? And now, let’s define that. What can be done to mitigate it or not? And who’s going to own that risk, at the end of the day? Is it going to be the chief legal officer? Is it going to be the chief security officer? Is it going to be the P& L owner? And ask them, ” Is that risk worth the trade- off?” Often, with mitigations, it is, or often, the juice isn’t worth the squeeze and you get to a decision and you can move on to the next thing. But I think that gray area of amorphous risk is really the enemy of progress.
Bradd Busick: And we can’t have risk without governance. So, I mean, it kind of is a throwback for me to the DevOps evolution, where we spent years as an industry looking at speed and stability as opposing forces before realizing that it was actually a system design issue, not an inherent feature issue of building software. Is that a fair parallel to where we’re at with AI governance right now?
Dr. Radha Plumb: Yeah. I like to quip, I guess, that the best analogy I have is that better brakes make faster trains, and this comes from the evolution of trains, where, of course, they counterintuitively were able to have the trains go faster in between stops when they had better and more reliable brakes. And I think of AI governance in that world, where, actually, AI governance is the fundamental thing that lets you know your AI solutions are doing the things you want them to do and not doing things you don’t want them to do. That’s it broken down. And actually, that’s also what you need for it to be effective. So to do something at speed, you want to build in those governance steps into the process, bake them in, and when you do that, you end up with a much more robust cake sooner, to extend the analogy.
Bradd Busick: When we think about businesses, they consist of people, processes, and technology. So with that maybe as the framework, and all businesses need AI, let’s set the stage for our audience. One of the things that makes LLMs so unique is even the model builders themselves still don’t have a full understanding of what’s happening in the, quote, unquote, ” little black box,” so to speak. So you and I could both have an identical prompt and create meaningfully different responses, and there’s nobody that hates that unpredictability more than an auditor, because their entire job is to verify why the system did what it did, and their job becomes incredibly more challenging with an LLM. So let’s dive into this can of worms. I mean, explain for me and for the audience how an LLM is fundamentally different from the type of systems that your auditor would normally be used to evaluating, and why you need to understand that gap.
Dr. Radha Plumb: It’s helpful, I think, to break the black box into chunks of where the box is. So there’s the inputs to that box, which are basically the data and the context. So by data, I mean literally the data, and that can be structured data, like your financial information and numbers, semi- structured data, like elements out of your contracts, or really unstructured data, like long documents or even images. And all of that goes into your algorithms with context, which is how is this data related to business and uses?
We’re used to combining those things and producing deterministic outcomes. So I take some context, I’ll use the most simple analogy, I take a bunch of data in a flat file, like a spreadsheet, and I apply a known statistical formula, like an average, and I put it in, out comes the average, and I can collect that many times over and get a distribution, or I can look at that over time and get a time series. Those are all deterministic outcomes.
What LLMs bring is taking that vastness of data and connections we both know and don’t know, and putting an inferential layer on it to come up with combinations of informations we don’t know and couldn’t have predicted to have a inferential outcome instead of a deterministic one. That’s the black box. That’s kind of the secret sauce. The benefit of that is it creates lots of things that you might not have been able to have before or even have thought of before. The downside of it is you don’t totally know the exact pieces that came to that or always how to recreate that.
So you really, I think, want to think about in your process, where do I want something creative, new, and different? And the LLMs are there. Where do I want deterministic outcomes? That’s where you can use your traditional analytic or MLOps, like traditional AI methods. They don’t all have to be LLMs. And then, how do I combine, what’s the control plane that mixes these together to produce the output I want that then is a predictable outcome for auditors, the benefit of generative where you need it, the predictability of deterministic where it has to be?
Bradd Busick: I love that lens. And if you think about this from a CIO’s perspective, who, in some cases, hasn’t gone down this journey, or in other cases, has gone down the journey, what do you anticipate the first compliance roadblock to be that a CIO should be worried about if they’re still early in that journey and about to move to an AI- driven system? What should they be thinking about?
Dr. Radha Plumb: Let me take it from the IBM perspective, just because it’s, I think, a helpful illustration. You have your data layer, and you need your data governance and controls. Oftentimes, for CIOs, that’s in the purview of a chief data officer, and there’ll be data governance and controls that you know. So the very first question you’re going to get asked is, how do you decide who gets access to which data that can get pulled in? What’s your role- based access control? What’s your identity and credential management?
So step one inside, for instance, IBM is we have a system that links your user ID, like most large enterprises, to your role and to that access. Now, you need to pull that data into your algorithmic system, and once you do that, you need something that orchestrates, is that data going into an LLM conversation? Is that data going into a deterministic poll? Is it going to just a dashboard? Is it going to a report? That control plane is something where you need the orchestration and the knowledge of the preferences of the organization, and that’s something you’re going to have to figure out and build for your organization. Enterprises are not going to be identical on this and it’s not going to be identical on different applications.
We, for instance, treat the orchestration of financial data very differently than we might treat the orchestration of rules related to brand colors and brand graphics that need to go into content. They’re both regulated. We can’t have 87 different kinds of blue for IBM. But we’re going to treat that differently than we’re going to treat earnings and revenue recognition in our financial systems, and there’s lots of stuff in between those two. So I think that orchestration layer is the first really big new question that CIOs need to start thinking about. What is your AI operating system? Where do you put that control plane, and how are you going to tailor it to your particular needs?
Bradd Busick: Yeah. I think that’s really well said. And when we think about what auditability actually looks like from an auditor’s lens, to be able to clearly articulate, to your point, role- based access control, this is what this individual has access to or doesn’t, this was the input and this was the output in that orchestrated control plane with a rhythm and an order and a discipline, easier said than done, as we both know really, really well. Because you’ve seen so many different types of these deployments in so many different industries, where do you feel like the biggest problem lies for most companies? Is it the data? Is it the model? Is it somewhere in the in between? Give me some color on that.
Dr. Radha Plumb: I would say it’s probably the intersection of technology and processes. It is a little bit both the data and model, but it’s really… There’s a feeling right now, I think, that you can take and sprinkle some AI magic into a process that is maybe too complex or maybe under- specified and that that’s going to drive really measurable business outcomes, and it’s just not. No CIO is going to be able to fix that on their own. So I think the real solution is trying to force the hard conversation about what the process should be, where the technology needs to integrate, what the technology needs to do, but where the process needs to change it well.
I’ll give you a concrete example. We’ve been working on this agentic workflow in finance to compare budget forecasts to actuals, a very common problem. And yeah, we can do that, and you need to think about the agent and you want to think about preferences on how big is a deviation that you’re going to care about. But also, you’ve got to kind of standardize the reports because you can’t automate variance detection for an infinity of cases. That’s not a technology problem. We can pick whatever threshold you want. That’s a process and controls question. And that needs to come in from the business and then be connected to the technology and that translation needs to happen. And then, that needs to be baked in in a way that is predictable and reviewable, so when our CFO says, ” Hey, why are we looking at this deviation and not that one?” There’s a clear business answer and then a technology solution that underpins that that can be demonstrated concretely to back it up. And that pairing, I think, is a big complexity that needs to get resolved.
Bradd Busick: It does feel like there’s a little bit of art and science here, and it feels like people conflate governing data with governing the model itself.
Dr. Radha Plumb: Oh, yes.
Bradd Busick: How are you thinking about this difference and in the work you’re doing and leading right now?
Dr. Radha Plumb: I try to think about this in layers, in some sense, because I think the data governance is, in some sense, a gating question before you get to any type of AI or digital solution. It’s like the fuel that’s going to power your AI model, so you’ve got to get that governance layer right. I think the problem often is you stop there. So you’ve got your data governance, you’ve got your metadata, you’ve got your role- based access control, you’ve got your authoritative systems, and you’re like, ” Great. Now, I’m going to AI this.” And you now have to think through, okay, once I’ve given this thing my data, what’s happening? What model governance do I need? What do I need to be able to see on what the model is doing, what data it’s accessing, the freshness of its data, how the model is performing over time, whether I’m seeing any particular biases? All of the normal things you would test in analytic solutions, let’s say, deterministic solutions, you need to bring back in.
But the problem we have is there aren’t known tests in the same way for these types of models. So what we’re trying to focus on is a lot more on the transparency side, understanding what the model is doing, what data is it accessing, when is it inferencing, trying to add transparency to the steps in the inference process, and use that to try to track where potential deviations and outcomes might come from. Hopefully, over time, we’ll also get better assessment tools, and the industry is continuing to develop those. But I think that’s really the complexity now, is there isn’t a baked- in, known way to test for accuracy or test for precision in the same ways we’re used to.
Bradd Busick: I think your call- out is spot on. I mean, it does feel like we’re flying the plane as we’re still trying to build it with regards to a regulatory environment that’s still being written. As you’ve looked across all of the organizations that you’ve had a chance to interface with, if an auditor came and sat down next to most of the CIOs you’ve engaged with today, do you think organizations actually have an answer ready to go when an auditor says to them, ” So tell me what’s going in and out of your LLM. How’s it being governed?”
Dr. Radha Plumb: It’s funny because I was just having a conversation with the chief investment officer at a large bank who was saying, basically, we’re not deploying AI on a lot of our investment decisions for this reason. We can’t be making decisions that we can’t back up what data is coming in and out.
There are things we can do and things we can’t do, and I think this is a good example of where we’re going to have to work with the auditors and with the regulatory community to create a reasonable middle ground. What you can do and what everyone should be doing is saying, for highly regulated decisions, this is the bounded set of data that can be used by the model to make this decision. And that’s going to, I think, be an excellent starting point for the regulators. I think then the second thing you want to do is using tools, again, we have one in IBM that’s called Watsonx Governance, but there are a range of these governance tools that actually tell you about what your model is doing. So where is your model inferencing? How is it behaving? That transparency is going to be really important for regulators.
And then, the last bit is when you’ve layered on agentic solutions… If you think about what an agent is, it’s RPA connected to a set of signals that come from a large set of data that you’re basically doing with natural language processing. So those set of actions need to be connected to the data. You ought to be able to transparently say, ” Here are the set of actions it takes. Here are the thresholds by which those actions get triggered.” And that now gives the auditor everything but the very granular way in which the data is transformed by the model to hit those thresholds, and I think that for most, but not all, regulated industries will get you through an auditor conversation.
It’s just worth the caveat to say there are some things right now we don’t have a good way to put into these LLM systems, and some of this is accepting the things you cannot change. So there are going to be things where you can do it and you can help use it to streamline, but it’s going to need to go into a human and a human’s going to need to review and make the decision based on the legal requirements, and that’s also a set of things that we should just not try to solve away with technology right now.
Bradd Busick: At this point, it’s kind of broadly accepted that security and speed are treated by most organizations like opposite ends of the seesaw. Speed goes up, security goes down. Security goes up, you’re slowing down my process. Do you think that this conception around speed and security actually applies to AI or can you actually have both?
Dr. Radha Plumb: I think you have to have both, and so you’ve got to flip the script on where security comes into the conversation. So a lot of our focus inside IBM, and this was the same at the Pentagon, is you’ve got to do security by design. The very first conversations I have with any new AI tool are with our CISO. I talk with him many, many times a day and I know basically all of his team by name. That’s not an accident. It’s because if I can’t get the security rules and insights they need right, then I can’t deploy.
And starting those conversations early so I know whether it’s a build or a buy decision, the questions they need to ask, the integrations they want to test, I know that on the front- end, and I can get quick answers and quick senses of whether this is going to sink or swim. That means that in the end, what we end up is an outcome that we know will be compliant and can scale. And that security by design, I think, is what lets us balance speed and security. And I wouldn’t even say balance it. I’d say creates this flywheel of security by design means compliant outcomes, means fast deployment, means you can do more security by design. That flywheel gets you going a lot faster.
Bradd Busick: I love your articulation of knowing the entire security team by name. I would say for our listeners, this is a really foreign concept. In some cases, the CISO is done to them versus with them. And yet, what I hear you saying is just the critical importance and maybe competitive advantage of having security and risk teams brought in at the beginning of an initiative or capability versus the end. Why do you think that that is so rare today, given the new world that we’re in?
Dr. Radha Plumb: I think a lot of times, people want to deploy solutions fast, and they think if they can just prove enough business value from the uses, that they’ll be able to bring the security team along. And oftentimes, that forces that risk conversation we were talking about at the beginning, where you say, ” There’s this big risk that’s going to cost us a lot to mitigate and there’s this big business value, P& L owner and CISO, who wants to hold which risk?” And you can have that and that is a way to resolve this, but it’s slow and it creates either risk or rejection.
We have found it way better to actually force a much smaller conversation at the beginning, which is, ” Where can we use this? How do we want to use it? What data are we going to use? What risks are we creating?” And with the CISO, a whole host of incremental mitigations and adjustments that happen as you’re doing your MVP build or your initial tests integration testing, depending on whether it’s a build or buy, you can do all of these things along that, that really mean that the end decision you have is, ” Hey, we’ve got a couple of risks here that we can’t mitigate. We don’t think they’re so big, given the business value. Let’s go.” Everyone feels really good about that decision. But that takes a lot more upfront work with the team, and people just haven’t mentally decided to move that whole process to the left. It’s a design feature. It’s not a compliance check.
Bradd Busick: Yeah. I love that. The notion of it being a design feature is spot on. I think it’s new, and I think for some, it’s foreign, particularly because they haven’t spent time thinking about how we actually want to plan. In some cases, they’ve been propelled right into, by the way, you have an AI platform, what are you going to do with it? So if you think about CIOs across the globe today that are sitting on platforms, that didn’t have AI 10 years ago, but had big data, and now they’re actually sitting on one that has agentic capabilities that are turning on nightly, what’s the one thing that you would tell them that they should start doing differently tomorrow?
Dr. Radha Plumb: It’s funny because it feels like it should be a technology thing and I’m going to totally talk about a process thing, which is the thing I would tell your CIOs to do is go understand the workflows and how they sit in your company. Again, I will use IBM, but we did the exact same thing at the Pentagon, which is we broke the business down into 10 big end- to- end enterprise workflows, and then in those, there are activity sets, and that’s how we think about deploying agents. But if you have that catalog, then as soon as you see new features, as soon as you see new capabilities, you can map those really quickly to the opportunity set on where they apply and how they apply, and bring those teams together and get a cross- functional team activating and deploying new technology.
But if you don’t have that upfront kind of boring process work that’s linked to your tech, so you know, okay, these are the different parts of our sales motion and this is how they’re linked to our Sales Cloud. Now, I’ve got my new features that were just launched to my Sales Cloud, or a new app that we’ve just partnered with and we’re buying, I should know exactly where those go and I should know who I can call up to say, ” Hey, let’s get a team together to look at this, do a quick 30- day test and see if it actually enhances productivity, and rinse and repeat.” And that’s kind of the approach we’ve taken, that that upfront investment is tedious, but really lets deployment happen at speed.
Bradd Busick: Well, I hope our listeners have had their hands free today because you’ve been dropping some knowledge. It’s been a pleasure chatting with you, Radha. Thank you so much for joining The Savvy CIO.
Dr. Radha Plumb: Thank you for having me.
Bradd Busick: I loved today’s conversation with Dr. Plumb. I think there was a couple of things that jumped out to me. One, bringing security in early and often tends to be the difference between success and failure in a deployment. And I loved her call- out of, ” I know all the security people by first name.” Imagine that at scale at a place like the Pentagon, where, I mean, honestly, you actually almost have to have that relationship in order to move things forward. So many CIOs that are listening today rely on their CISOs and their security team for all the things that nobody cares about until they break.
I think Dr. Plumb’s call- out on understanding workflows is everything. In the absence of workflows, you’re going to throw AI at something and hope that something awesome happens, and as we all know, businesses don’t run on hope. So I think spending time to apply the discipline, to documenting your workflows, understanding that so that during an audit, you can marry up the workflow with the technology is a better recipe for success.
That’s all for today. Thanks so much for listening. Please follow us so you don’t miss an episode. This has been The Savvy CIO, brought to you by Park Place Technologies. If you want to learn more about Park Place, go to www. parkplacetechnologies.com. And now, one last word from our guest. Radha, because the show is called The Savvy CIO, what is the savviest choice you’ve made in your career to date?
Dr. Radha Plumb: I think it was deciding to go all in on figuring out enterprise AI. I think it’s going to be the place where people are going to spend the next five to 10 years really transforming everything about society, and it’s really exciting to get to be a part of the story.
Bradd Busick: I love that take and I couldn’t agree with you more. I’m your host, Bradd Busick. And as always, IT shouldn’t just be at the table. IT is the table. Be well.