
AI Engineering Beyond the Hype: Harnesses, Guardrails and Production Reality
WHAT MAKES AN ENTERPRISE AI SYSTEM PRODUCTION-READY?
A production-ready enterprise AI system is observable, testable, cost-controlled and governed throughout its lifecycle. It needs explicit success criteria, representative evaluations, runtime guardrails, secure data access, version control, incident handling and clear ownership—not only a capable model or an impressive prototype.
KEY TAKEAWAYS
• AI engineering adds probabilistic behaviour, context, evaluation and runtime-control concerns to software engineering.
• A demonstration is not production evidence; reliability, security, operating cost and failure handling must be tested.
• Engineering harnesses should apply policy, testing, telemetry and release controls consistently.
• Context, retrieval and enterprise memory need lifecycle, access and quality controls.
• Model choice is only one component of a governed production system.
SOURCES AND FURTHER READING
NIST AI Risk Management Framework: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10
Australian Government AI technical standard: https://www.digital.gov.au/policy/ai/AI-technical-standard
OWASP Top 10 for Large Language Model Applications: https://owasp.org/www-project-top-10-for-large-language-model-applications/
RELATED EPISODES
The Agentic Enterprise: https://www.enterprisetechtalk.com/episodes/agentic-enterprise-scaling-ai-controlling-costs-redesigning-work
Agentic Governance: https://www.enterprisetechtalk.com/episodes/agentic-governance-enterprise-ai
From Systems to Agents: https://www.enterprisetechtalk.com/episodes/systems-to-agents-enterprise-architecture-aiGenerative AI has moved rapidly from curiosity to enterprise priority. Many organisations are now exploring AI-assisted software delivery, coding agents and agentic workflows. But as the technology matures, one reality is becoming clear: building an impressive prototype is very different from engineering a secure, reliable and cost-effective AI system in production.
That is the focus of this Enterprise Tech Talk episode, “AI Engineering Beyond the Hype: Harnesses, Guardrails and Production Reality,” hosted by Saumitra Kalikar with guest Adam Witanowski.
Adam brings around three decades of experience across software engineering, startups, consulting, enterprise technology and AI architecture. His perspective is grounded in practical delivery: AI is reshaping engineering, but it does not remove the need for strong architecture, testing, governance, security and disciplined software delivery.
FULL TRANSCRIPT
This transcript is based on the episode’s English auto-captions and has been formatted for readability. Please allow for occasional transcription errors in names, acronyms and specialised terms.
[00:00:00]
Atlassian study engineers only spend 16% of their day doing um actual coding right and so uh we're really optimized at the moment for that 16% where we're filling with characters um and and that was never the job of engineering. I think a lot of organizations are stuck in the conversation around which model they're using. Um, which is a very, you know, 2025 discussion. Um, 2026 discussion needs to be around, you know, how are we empowering people with better harnesses and empowering the whole world with the meta harness that we're building. Just like you, you know, a human needs to uh needs to have um some framework around what it's building. This is what specdriven development obviously uh you know gives it and so um you know some people see specri development as you know just the PRD where it's you know a definition of of uh various pieces and I find that a little bit loose now nowadays QA is I think the the biggest sort of unsolved problem at the moment in uh in the SDLC agentic SDLC uh and I think it I think it falls it's a problem in two ways one is we've always talked about QA shifting left.
[00:01:15]
Um, and uh, the faster we move with with agent development, the less likely we are to actually shift the QA left. A lot of companies, their their appetite is going to be to jump straight into building solutions uh, rather than um, first of all wrestling with some of the ethical and some of the um, you know, and some of the regulatory things that they they really need to cover off first. Um, and what what they'll find is uh that'll become just a massive break, like a a massive uh break for them, unless they sort that out first. Hello and welcome to the Enterprise Tech Talk podcast. I am your host Somitra Kalikar. Now AI has moved rapidly from just being a curiosity to an enterprise priority. Many organizations are now actively exploring uh assisted software disco delivery and uh agentic workflows.
[00:02:26]
uh organizations are now discovering that um the building a flashy AI prototypes is vastly different than engineering secure, reliable and cost effective AI systems and that is why today's conversation focuses on AI engineering. It's not only about prom engineering or selection of new models but the broader discipline about how to build reliable AI systems in production. And to help me unpack this topic, I'm joined today by Adam Vidonoski. Adam has a wealth of experience at the intersection of AI software delivery and uh enterprise technologies and with his insights. Today we will explore how to move from a prototype hype to a production reality. Adam, welcome to the podcast. Great to have you here. >> Yeah, great to be here. Thanks for having me. Um, Adam, before we get into details, uh, tell us a little bit about your professional background and how you got into AI engineering in the first place.
[00:03:31]
>> Sure. Um, well, uh, I've been a been in software for about 30 years. So, I've been a builder. Uh, I've run my own companies, done startups, and and predominantly around that um sort of startup and and, uh, consulting kind of space. Um, NIB uh is where I was at last as a uh AI architect and prior to that uh I was in the same organization as a senior engineering manager um working with a you know a relatively large team um in the uh in the travel insurance domain moved across into the group domain for AI architect. uh obviously when AI started you know it sort of entered entered the uh the common man's world through chat GPT you know four four five years ago uh this um little chat box that we that uh it just it just seems like such a you know simple and almost pathetic way to interface with a profound technology and yet this is what we've landed on um is uh and sort of entered into the world and uh and immediate mediately I think like everyone um was just blown away by the potential of it. And so of course uh I started you know digging into it and and trying to understand how this new attention is everything idea works um looking at how you know the the underlying models work and and vectorization etc. and um and became quite enamored with that as an idea and started to sort of steer myself more and more towards that. As a senior manager, you tend to not be on the tools that much and I found myself on the tools every night. Uh you know, >> playing with it, understanding with it, using it for development and as APIs and whatnot became available, um building products with it. And so uh yeah, I've been playing with it for um you know, four or five years. Obviously following it along, um various other iterations of uh of ML prior to that. Um but particularly in my last role for the last 12 months, um running functionally a lab within uh within NIB doing experimentation and ultimately product delivery uh with uh with AI. Um so it's been a it's been like everyone a wild ride. changes every five minutes. Uh you have to be reading the news every two to keep up. And um it's it's it's been a lot of fun obviously.
[00:06:03]
>> Yeah. And as you said uh we all got exposure to LM based AI around four years back. Um uh but in the last one year or so I would say the the the AI engineering the uh that discipline has started becoming much more mature. Right. Um and do you see do you think we are at a stage in terms of maturity of uh AI based engineering and coding that it is now becoming its own distinct discipline compared to traditional way of engineering? >> Um yes and no. So the the the concepts are the same and so like I think if you were in an organization uh you know humans make mistakes all the time and uh and we make coding mistakes all the time and and we're you know perfectly for people all of us. >> Yeah. Um and uh if you're an assistant, if you're in a organization that has built good guard rails around uh around human engineers, then you're probably already directionally correct towards having uh you know an organization that's reasonably well or at least directionally set up for having um agents um also write code because you know good good environments are uh are environments that are intentionally low trust. So you know they we have a good PR discipline good CI/CD uh pipelines um ways of doing um you know validations on data that's that's coming in and data that's going out all of those are still required disciplines uh for an organization to have with AI u typically and this was a this was a Atlassian study engineers only spend 16% of their day doing um actual coding right So, uh, we're really optimized at the moment for that 16% where we're filling with characters. Um, and and that was never the job of engineering. The job of engineering has always been problem solving, architecting, um, and imagining new ways of solving problems. And so, I I think that that as a discipline is still very much at play. uh and uh and the you know optimizing around that 16% and the various other things that you need to build around that to make that um safe and and secure for deploying into an environment. um you know functionally ex exist but they need to be scaled very differently because you know if you can imagine if you've got 16% of an engineers's time you've got an organization with you know 100 a thousand engineers and um and all of a sudden that's through that same pipe you're now shoving 10x of the the code through obviously your PRs are going to break your QA processes are going to break uh and your um you're shoving things into that pipeline through um product managers, product owners, BAS, uh and their back end as well, your cyber security uh and your cloud.
[00:09:09]
They're going to start to strain under that. So, you do need to think about how do we widen the whole pipe, but um but functionally, I think we've we've thought through these problems um historically. We just need to figure out how to scale them in the same in the same way we've scaled uh the code generation. Yeah, I think uh it's wise to say we are building on top of all the fundamentals that we have built as part of the standard software engineering practices. Uh but that 16% that you said that is unique potentially uh that definitely I want to unpack in terms of what it means when it comes to the designing the reliable and predictable AI solutions etc. But before we go there, this space is definitely is filled with many hypes, right? So I just wanted to pick up your thoughts on at a high level which of the trends you see are um overhyped today in the AI engineering space and and more importantly which trends are have emerged but are underrated which are not talked about that much but should should have get that much more attention.
[00:10:19]
Yeah. Um, you know, you go on LinkedIn and, uh, every every man and his dog is on there saying, "I built Spotify over the weekend." >> And, um, that is, uh, good for you, brother. Guarantee you didn't build Spotify over the weekend. Uh, you built something that was perhaps a nice clickable UI. uh and if you try to launch that into production um you are going to you know you're going to lose your customers data data money etc. Um so they took pock and um and AI is incredibly good at building pocks. Like just this morning, I I I've had an idea for uh for a while that I've been going back and forth on. It's quite a complex um problem space to solve in and I thought I'll just I'll just start, you know, work start from the UI and work back towards the back end that I've started to build out. Uh, and I got a I got stunning incredible demo, you know, out of a PRD uh that that I've spent, you know, months um re refining, but um, you know, it it is it's not a product, you know, it's so far from a product.
[00:11:27]
Um, so there's definitely this this hype um because we're and it's it's addictive honestly like this, you know, how fast you can spin something out that is that is functional, clickable uh and uh and can solve a real problem. And in some cases that's enough. You know, if if you're solving a if you're solving a problem just for yourself, um you know, sometimes that is that that's completely sufficient. Um but if if you're going into uh a production environment uh you need you need far more infrastructure around that process uh than uh than you know I built Spotify over the weekend. Um and you know the other I think the other hype piece is that you know it's going to replace engineers in the next 3 months six months. I I don't believe that for for a moment. I I think eventually over time perhaps within the next two years it will I mean it's already reshaping um organizations and how they think about uh how they think about engineering and other disciplines within the organization um and over time it will dramatically change like without a doubt the the the um the the discipline um but not today probably and enterprises are typically slower to adopt and so you know that that transition take a while. Um but the the underhyped things I think are around the importance of um of harness uh and harness engineering uh and um and something which people aren't talking about as much uh two things. one is multiplayer uh aspects of of AI and your AI use cases um where we you know at the moment we've so we've spent the last couple of years training agents how to work as a team um how to uh how to swarm over uh over problems um and we've we've left humans out of that as this as this single player uh you know in through a through a chat interface um whereas I think we need to broaden that out into a much more multiplayer mindset where we incorporate other parts of the organization, other parts of engineering to be able to influence the outcome. Um, and people tend to talk about it as harness. I I talk about it as I met a harness or a harness harness because it's the things that outside of you as a personal developer working on something to start to include the whole team um to be able to shape just the influence that's happening on your machine. So um that's where I think we just need to start moving towards a conversation of multiplayer and and and uh and meta harness or harness harness. Uh but certainly at at an individual level um that you know that the harness is is an incredibly important idea. I think a lot of organizations are stuck in the conversation around which model they're using. um which is a very you know 2025 discussion uh 2026 discussion needs to be around you know how are we empowering people with better harnesses and empowering the whole or all with the meta harness that we built.
[00:14:42]
>> Exactly. And that was I wanted to unpack that anyway. So good you actually given the direction on that. Um um the if you had to take a deep dive the first point I wanted to unpack with you was about how we ensure reliability of u uh the AISS systems uh in production and uh as you rightly said um selecting model is just one fraction one one part of it um but that's not everything the more important when it comes to reliability aspects in particular is how you ensure the security how do you ensure the the the the overall reliability of of the system and building that harness around that becomes incredibly important and industry has shifted from the product the the prompt engineering to now focusing more on harness engineering. Right. So if would you mind unpacking that when when you talked about harness what really goes underneath what are the key control planes uh organization should be thinking of to build a good harness. >> Yeah. and and you know uh for when people talk about harness they're talking about multiple things and so it's just a it's a it's a catch all bucket at the moment and I think we need to start to distinguish um what we mean when we say harness because the harness could be um you know the the various tools and uh and memory types and things that you use in an agent >> um it can be you know clawed code and uh and uh codecs uh and um and cursor are all examples of harnesses in their own right. Uh but then when engineers start to talk about harness, they talk about the custom tooling they they build the sub aents um they get called during invocation. Uh and so there's that uh you know there's those ideas and then I've also seen people talking about harness in terms of the whole team um infrastructure that sits around uh you know the single individual engineers. Um when I when I talk about uh you know harness or I I typically talk about a meta harness because I like to think in terms of whole teams um an individual is going to build the way they build right and if they some of vs code or jet brains historically you know they pick their IDE um they you know engineers are going to build their their own coding harness uh setup and that's going to become their their secret source their magic you know like like a Jedi building a lightsaber uh and and it's going to be unique to them. But the magic wasn't ever, you know, if you look at Star Wars as an idea, uh you know, we we're enamored with the Jedi and their lightsabers. Uh and uh but that's just a that's just a tube with a with a crystal in it and some force magic. you know there is probably a bad example but you know the Death Star now that is a >> that is an architectural marvel you know and it's a what a stunning picture of civilization level uh collaboration and and so I think really the power going forward is going to be at that's you know at that sort of meta scale uh where we where we start to think about harness harnesses or meta harnesses um so you know some of the things you you mentioned cyber security um there's obviously things like QA say uh there's um there's cloud infrastructure marketing has a say legal risk compliance they all have a say uh in how things get built um and and uh models and coding agents are really good I mean really good right now you know since December and you you had um what was it Opus 4.7 uh come out around December uh they they've been they've been astronomically good at generating Um and and since you know you've had uh various GVT models come out and they're also exceptionally good. Now you've got Chinese models GLM 5.2 and uh and Kim K 2.6 all very good at generating uh code and so like personally I haven't I've barely written a lot of code in 12 months you know like I've written code of course but um it's it's probably down on that 10%.
[00:19:00]
Um, and when you're doing green fields, that's totally fine. Like if I'm if I just put in a put in a PRD or a prompt or something and have it generate some code, uh, and I don't care that much around, um, you know, style or language or, uh, infrastructure or databases, it's going to turn out something that feels like black magic. And that's, you know, that's >> every man and his dog on on LinkedIn who built Spotify. Uh, they're able to do that. uh but in terms of generating code particular to your organization uh that's a that's a completely different story. They don't know the the model doesn't know anything about your organization. And so this is where we we need to shape a meta harness around it to shape the outcome together uh that that's unique to the to the needs of your organization. And so you know one one thing that I uh built uh was it was called the governor module. Um and that allowed actually two things that I'll I'll I'll reference. One was governor module. Uh and that allowed um vertical teams. So uh teams like cloud cider um you know data every every engineer always screws data over because we you know we make changes to the model that we don't tell them. So they're adjusting you know broken pipelines etc.
[00:20:18]
um risk, compliance, legal, marketing, they can all define their policies uh and those get converted to policies of the code uh and injected um at invocation time into the coding agent either via hook or MCP. Um similarly, horizontal teams that are doing the actual can define their uh their um particular teams um style guides etc. and that gets injected at invocation time. And so now you have a control plane that sits over over everyone's coding agents that um that that's uh that that is getting injected at invocation time um the various policies of various parts of the organization and and this is what I call giving agents good steer um the engineer can do good steer for sure uh but when time crunch happens they they're probably not thinking um about the you know 150 page documents that uh is sitting there in confluence that define our cyber security mandates and and how we and the agent definitely is thinking about that.
[00:21:26]
Um but if we define policy as code and injected an invocation all of a sudden that shapes what gets delivered to be more organizational shaped and that can you know include things like how we deploy how you know what cloud infrastructure we prefer how we like to secure APIs etc. then um and then since you have that policy as code, it's very easy to then reassert it at CI and so uh you can you can virtually guarantee that um that what you know uh it needs to be the shape of code for your organization ends up getting delivered in a fairly similar shape to what you need for your organization because you've uh you've worked collaborative collaboratively as a team and if that changes over time or each of those teams can own that vertical injection. The other piece is um coding agents are very good at working on the individual codebase that you are reading. Um and so if you open up, you know, a GitHub repo and it it it looks at it and it's able to reason over it and that's great. The problem is organizations very rarely have a single codebase. Um NIB had 3,600 plus repos.
[00:22:37]
uh and you make a change in one of those repos, there's a high potential that there's going to be an upstream or downstream impact uh you know to at least data or another consuming service. Um and so we need a way to give visibility to the coding agent for uh repos that sit outside the the the current thing you're working on which you may not not even know as an engineer you're making downstream. Um and so uh something that I built uh which I called Skynet rather ominously ingested all the repos into a knowledge graph uh created a vectorization uh and um and extracted code intelligence um uh code intelligence database out of that as well. And so gave the agents uh the coding agents of the engineer a 360 degree view of all the other repos uh that it could pre-calculate um downstream impact and upstream consumption um before it made a change. And so, you know, neither of those are are cheap um things to build and neither and and there are some examples of these like source graph um has something like that for the for the Skynet.
[00:23:47]
Um but organizations need to start thinking about how they shape that meta harness to create good steer for the coding to make organizational shaped code, not just code in general. Yeah, clearly you are a huge fan of Star Wars and Terminator. So, uh we we are definitely aligned there. Um look uh and I think uh one thing you said uh and it's I mean if I had to summarize is traditional soft in the traditional software engineering most of those control planes in a in a way where more reactive in the sense that you build a code then someone will go through a a security governance review. you'll go through any of the different types of reviews but that is typically much more reactive um uh u in as in the AI world when you are talking about the 10x engineer uh uh workload um and uh the the work is getting accelerated you need to shift from reactive to more runtime controls right and then that's where that harness becomes important that it has to be more runtime um than being reactive uh But my question also is given given this is the emerging capability need uh do you see industry solutions emerging as best of breed those runtime harnesses um or do you see organizations still need to build their own harness models?
[00:25:18]
>> Yeah, it's um it's fragmented. Uh so you know going back 12 months when I started the AI architect role there just wasn't a lot of these solutions. You sat to you had to roll your own and they fit into a category of what I call boring problems. Um and so you know uh doing doing um PRs is a boring problem. doing agentic PRs now is becoming increasingly a solved problem because it's it's the clear next bottleneck after shipping right is you're you're gonna uh you're going to ship a large volume of code and it needs to be PR and historically uh those have been poor anyway right you know we we tend to in engineering rubber stamp uh stuff you know looks good to me uh approve um which is a bad way to do PRs but it's probably how a lot of PRs happen in companies anyway. Um, you know, and so the the agentic PRs uh and there's some debate about how best to do this and uh you know, should it be the same model as doing the invocation and you know, whatever you we can debate that um till the cows come home, but uh at the velocity that we're that we're going to see code getting shipped um that's 100% need to is going to need to be authentic. Um, and so in some s slices of of the uh of the um SDLC, we're starting to see these tools uh emerge. Um, and I and I think that's, you know, that's great and and and but companies are already starting to cry foul at the cost of tokens.
[00:26:53]
>> Um, these are these are just more token consuming machines, right? Right. And so if we're not comfortable with the invocation tokens, we're probably not going to be comfortable with PR tokens, with the cyber tokens, with the QA tokens. Um, and so I think we need to, um, we need to either get comfortable with, you know, with token spend or figure out better ways to squash the cost of tokens or get more comfortable using, you know, open source models, uh, and model variable model um, uh, variable capable models. So, you know, throwing out some bits to say Hi coup or Gemini um 3.5 Flash for cheaper operations uh rather than throw everything at Fable or or >> Okay. >> Uh and that's typically, you know, not not handled that well in organizations. Everyone's throwing stuff out to superm models as opposed to, you know, figuring out these nice thin slices.
[00:27:51]
>> Yeah. Let's let's shift our focus on um uh on the SDLC the software delivery life cycle approaches emerging in in AI engineering and and when we started AI based coding as you know um we started with just asking uh AI models chd or sorry codex or um uh claw to build a small feature build a small tute a small task and what anthropopathy that time coined pipe coding and um uh we and that's the uh term that has stuck with us. Um but as we have mature the AI engineering practice um the shift has uh uh has been around uh not just giving a task but doing a proper uh spec based development specular development so the so that um uh the AI doesn't do the coding just by intuition but it has a very clear view on what is the requirements what is the inputs and how it should test the outputs etc. And now we are also showing the shift towards more goal-driven development. Um right which is you don't necessarily go and intervene at every stage but you define design a very clear goal and and then let AI iterate towards achieving that goal. How do you see this these shifts emerging and particularly I want to get your thoughts on when we talk about specdriven development what best practices are evolved there and how spec development and goal-driven development which is now being taught more nowadays how they're going to complement with each other >> yeah sure so I I I before it had a name I was doing spectrum development I my first my first use of AI for coding was copy and paste like a lot of people right you know you ask a question and you know basically using using chatb like stack overflow or something is spitting out an you copy and paste that into your IDE and run um you know obviously that was inefficient and so uh companies like cursor and uh and claude code uh saw that you know single line complete and and kudos to GitHub copilot you know they were very early on this uh and they were doing this single line completion um and uh And you know that was great uh for for what it was. Then the others thought what if it could just complete more than one line? What if it could complete a whole file? What if it could reason over multiple files? And so cursor and cloud code come in and start to do that uh and say do the copy and paste as well larger chunks. Um obviously as you start to handle more or hand off more to the coding agent uh it needs more context. And so uh you know if if you ask it uh you know build me Spotify um it has a it has a really great reference to go from it can it understands what Spotify is. It understands the branding. You can probably go off and search it and find you know the uh the DOM and uh and return you something that's pretty similar. Uh in most cases you're not going to be rebuilding Spotify. you're going to be building something particular to your organization, uh, a feature that's particular to, uh, some functionality that exists in your software. And so its references are going to be, you know, more limited. And so giving it just like you, you know, a human needs to, uh, needs to have, um, some framework around what it's building. This is what specdriven development obviously uh, you know, gives it. And so um you know some people see spectrum development as just the PRD whereas you know a definition of of various pieces and I find that a little bit loose now nowadays um whereas I like to use um in that context I like to use uh because I also like testing um I've been experimenting more and more with say girken and and uh and cucumber style practices where you give it uh you know smaller smaller chunks which become ends up becoming tasks Um, but I also like to then go through and do an architectural piece and a uh and a solution design piece. And so I I can get more granular in terms of the steering, the upfront steering of what's going to go into the coding agent. Now um and so but I would consider spec uh to be you know really clearly defined um uh you know thing that's going to be built. So a product requirements document um with some breakdown of um of the of the end state uh user you know usage requirements and and what basically good product practice uh and then um and incorporate some early injection of u of QA practice using uh you know things like cucumber and girkin style language um and then move into uh the architecture and so I'll have a separate architectural document um where I define um you know it might be uh I get the original architecture of a piece and then extend that uh and you know that gives me some assurance that it's going to land near where I want it to land uh and then a solution design dock similar to what I do and and honestly the models are very good at generating almost all of that as well. So your spec >> generate lots of it, review it, and at least you have reviewed.
[00:33:14]
Um, anyone who's been doing specri development for a while knows that the practice ends up becoming, okay, I give this to Claude. Uh, Claude, you know, does and I always do, you know, a phase delivery. So give me, you know, I want this to be a phase delivery, multiple tasks per phase. Um and uh you know and and it will split out that those phases into tasks and then quite happily tick through those tasks. Um you know what happened about six or seven months ago people started talking about loops uh and you know this obviously the Ralph Wigum loop and others um uh things like uh Gas Town um which automate you know large portions of this process. GSD is another one. Um and uh and so the what you what they found was or what engineers were finding was that the model would go through or their cord would go through as an example um and build task one, task two, task three and it's like are you ready for task four?
[00:34:13]
And so their their next action was yes enter you know and so their involvement in the development process and ended up becoming saying yes and hitting enter. uh and uh and hopefully they had some other you know um skills and things in there or some uh something in their claw MD or what have you. Uh they were steering the agent to also build tests and other things. Um and so you know their their evolvement became less and less the better the models became. And so then you start to go well why bother hitting yes enter at all? Why wouldn't we just loop through that um and get you know through full completion? Um now what I found with what I found with a lot of the sort of uh you know the particularly the early loop test and now there's now there's um you know some more defined golf um earlier models would hit about a 70% good outcome. Um, now we're getting to more like an a probably an 80 to 85% good outcome. And you still need to do a lot of, you know, uh, rework and testing and and but most of it I find now at the end of the loop or or at the end of a goal is I'm educating myself on what was built and getting it to replay to me and prove to me that what I've built was good. Um, but these, you know, I've had loops run for for just, you know, hours, whole days. uh and at the end of it a I'm poorer uh because the token costs uh but b um have a just a huge corpus of code that now I have to try and understand um but functionally uh more and more and particularly with you know models like fable and and gtd 5.6 six. Um I'm seeing that there's there's code that is that that is just very good or or outcomes that are very >> and and the sorry no but uh so the the loop engineering is definitely becoming uh attractive and uh there's a lot of optic to that but there are also a couple of concerns and and one you just mentioned is um if not properly designed it can lead to token maximization. Um that's one thing and second thing is uh as you keep uh as AI keeps running into loops and the context keeps piling up building up um uh the more context that is built there is always um a risk of drift u and um how so so what are the best practices that can are emerging or what are the guard rates should be put in place in designing the loop so that these concerns are are addressed.
[00:36:52]
>> Yeah. So um one just I mean this is a really simple one a lot of people already doing it is um is you know first of all sub agents obviously in your loops and how you define those and um and so you know sub agents will obviously get their own context and so uh so heavily heavily relying on sub agents um but then having a I like to use a good model as like the orchestrator model that sits over the top something you know now fable obviously is a good choice is to define all of the tasks and define all of the um uh you know the the what's ultimately going to get built. Um hands off to smaller models like Sonnet. uh Sonnet 5 is very capable um at generating code particularly if it's going to be validated by a larger model like Opus or or Sonnet um and uh and in some cases you can have um even Haiku uh do do things like documentation and whatnot which are also uh by volume expensive exercises right uh and uh um and so writing tests or running tests or u or running the code and watching the output haiku's great for that. Um, and so I think you know the the cost savings I mean there are a couple of a couple of ways. One is um in organizations this is a little bit harder because your cash pool is often smaller if you're running out of like an enterprise bedrock environment um but but trying to optimize for c hits and and and that sort of thing. squashing um squashing your uh your inference uh inputs and outputs with something like tune or JSON reduction in some way. Uh you know we've seen 30% plus reductions um there. Um and so you can minimize the the raw token burn, but then uh and the the raw query burn um cost you can you can definitely do by by using um sub agents that you clearly define um and uh and hand off a lot of tasks to smaller cheaper models. Um but I typically like to have a larger model driving um and then in post validation uh doing the doing the code review.
[00:39:06]
>> Yeah. Look um the other other uh uh point I wanted to unpack with you was around u uh the the testing uh and the evals for for testing. Um now in in the traditional uh software engineering and testing the traditional safety systems is all about you have clearly defined test cases test criteria you run those test cases after feature is built and um and you you either pass or fail essentially the test case right that's not as binary um in in the AI AI AI driven AI systems right um how what what are your um uh recommendation in terms of building good test practice to support um AI based developments. What are the runtime events which um engineers should be thinking of? Yeah, I think you know out of the full SDLC you know uh cyber used to be the one that worried me the most uh because obviously you know there's high risk there uh but but you know companies like whiz and and others have come in and they've built really great tool sets around that space and and there's continuing work happening there and and and of course um you know big companies have seen the risk and AWS is releasing stuff Microsoft Google and Um and and so that's less of a that's of an unsolved problem. QA is I think the the biggest sort of unsolved problem at the moment in uh in the SDLC agentic SDLC. Uh and I think it I think it falls it's a problem in two ways. One is we've always talked about QA shifting left. Um and uh the faster we move with with agentic development, the less likely we are to actually shift the QA left, right? uh we we actually because the engineer can just run ahead and and they often don't talk to the um to the QAs on the team um around how this is going to be tested you know what are the what are the uh you know what's being built and often the engineer doesn't even know that right they might have an idea from the spec but they don't know exactly the shape of it um until it's built um and particularly in an environment where you're uh giving less steer up front um so I think um you know in in terms of teams and enterprise teams, they they if there if if there's ever a time for shift left uh for for quality assurance, it's now. And so that needs to be quality assurance over the spec, it needs to be quality assurance over uh over the meta harness and understanding how that's going to deliver what it delivers. Um but more and more the the job of a QA is going to be less around um is less around testing and more thinking like an engineer and acting like an engineer. Uh and uh definitely at NIB we saw we're seeing the best QAS uh who were keeping pace with the um with the with the AI engineers were the ones that were thinking and acting bike engineers themselves and building um their own harnesses around uh around what's getting built. Obviously you got things like golden eval golden um golden data sets to eval against uh you know uh um building um of course most most of the time models won't volunteer to build u unit tests and u integration tests and um and so I think measuring that the test volume of that certainly um you know playright or or other endtoend uh testing um frameworks and practices are increasingly important and co Claude I'm sorry Anthropic has said multiple times um giving the agent the ability to test its own output is really important um and often engineers don't include that in their harness and I think that is that's a that's a real problem um definitely uh involving the QA very early uh and increasingly they're going to become engineers in of their own right building their own um or and this has always been true. There should be a tension between QA and engineers, right?
[00:43:21]
It almost um uh like the the QA their job should try and be to just fundamentally break anything the engineers build. Um and uh and I think their ability to do that they're I mean they're becoming just so outgunned uh with what you both the pace and and um and the visibility that that's more and more they really are going to have to be engineers who deeply understand the spec um and are able to build frameworks at the same pace as the engineers of building code. >> Yeah. Cool. Um uh the other um uh point I wanted to unpack with you Adam was around um the context management and the memory um enterprise memory management you uh and as you know you and I know uh that this is a unique thing about agentic systems that they are more effective um when the context is right u right uh that is not something that was uh relevant for traditional deterministics software. Now we um the the so the the industry has evolved on that as well in the sense that um we started with uh designing um the retrieval augmented model model applying retrieval augmented models to get the data organized context into uh the software those systems. Um but now we are also seeing um the emergence of large and availability of large context windows um for this uh agents right. Um where do you sit in in the in this debate uh whether rag is more effective or large context models are is is the way to go or there is something um they need to complement with each other.
[00:45:16]
Yeah, I think it's I think it's um actually a third option. Uh so, you know, rag rag is great if you want to uh if you want to understand does something exist, you know, in um in a document. Uh and um and but it doesn't give you why it's there. It doesn't give you um who you know made a decision. Um and you know there's of course things like you know I think PRD become if the repo becomes the source of truth PRD should be shipped with the repo uh and um and that becomes you know something like the why a change was made now managing that over time becomes complex right because you you in one one way of doing it I've seen people do is they just you know they update the PD to match the uh the um the code that's shipped and so you have a record in the git in the git history of changes over time um and why things exist where they exist. That's a you know and then the other side is people store feature um you know feature level PRDS uh and so they have a history of change over time um but the the context that that really matters and missing often from you know enterprises now that aren't using agentic delivery and it exists usually only in the head of a couple of senior engineers is why the codebase looks the way it does. uh why was change made um you know 14 months ago that's now causing us this problem in production I mean you have to go and ask this guy that's been there for 15 years uh the problem is these people leave and so that knowledge leaves with them and so I think there's an opportunity here with uh with you know agentic delivery to start to think about this problem again as a harmless problem or as a meta meta harness problem is capturing um the intent not just the you know uh the the change that we're going to make but the intent who who made it who made that decision why that decision was made um the point in time when that became a a a depre a deprecated idea um because you know like human memory we we learn new things and and we degrade old memories over time um I think uh so the third option I think is some combination of larger models for sure but you know I think that that's going to threshold um and you see that already like in a you know in a million token context window um you're really only getting you know sort of 500 um thousand tokens worth of useful uh memory there and after that it starts to degrade quite a lot in fact it starts to degrade you know after about a quarter the way through we start to see it you know slip slip a bit um and so I think we need to refine what gets passed in rag does that in part but as I said it doesn't give you any of the any of the the why or any of the when or any of the um you know the decision ledger or change in decision over time. So I think we actually need to have some combination of rag and graph um where the graph is uh is um the nodes are keeping track of um decisions over time um and the specific parts of uh of you know because you rag a you rag a a document or rag you know memories um and uh and you're going to return this tons of results uh and those may or may not be the context that you want. Um and so then you're again throwing that into the model to decide whereas a graph is going to be a little bit more deterministic around the choices that are made. And so you're starting to see companies like me zero and and others um enter into this space where they're uh where they're building um graphbacked uh you know memory systems and I think more and more this is going to become something which organizations want and need. Not only are they going to have their business context, but they're going to have this code context >> uh where you know we're going to be capturing the episodic memories out of every coding agent uh and and starting to store that in some way uh and referenceable through a graph so that every my coding agent has the same context as your coding agent the changes that were made over time. Uh obviously that's not really in place yet. You make a change in the code, it exists in your context. um and I don't have it and as soon as you start a new context um it's lost forever right and so we don't have any history of the decisions that were made we don't have any you know why that decision was made who made that decision and these are actually really important for organizations to have not just for you know for reasons for blame which we you know humanistically we like to do uh but for for reasons of um of context particularly as we get to you know uh which we haven't even spoken about yet which is you know regulatory and compliance problems particularly in the FSI space um you know where where is the human being in this process uh who drove these decisions why is the code shaped this way uh and so we're left being archaeologists of code without really having a map to be able to go through the history of the changes and so I think yeah graph graphback memory um where where we start to record sort of that episodic movement of of of invocation through through coding agents and uh agents running in the cloud etc.
[00:50:46]
um is going to become more and more important over time. Um so yeah I I think it's a combination of both but a third piece which I think is being able to actually point and time stamp it um with the decision and who made the decision. Yeah. And as organiz organizations are at the early stage of building designing and building these um context complaints and enterprise memory layers and and as you mentioned episodic is one of the layers etc. But I see as organizations start building more and more agentic systems um they will have to start considering on uh about how to optimize the context. For example, because >> um uh how to cache memory um because uh u different agentic workloads may require different context and just enough context that is also important because too much of context may also drift um the agent. So how do you see that those emerging concerns may not be the concerns today for organizations because that's so early stage but as this maturing these will become real concerns for them and how they should be thinking of addressing. Yeah, definitely. I already think I already think a lot of engineers just as an individual level over stuff um you know over stuff their coding agents and and I think there's always a danger that people want to over stuff their you know their cloud workflows as well like agents running better etc. Uh the reality is the reason it's not a problem for a lot of organizations right now is that they're that their data in general isn't ready. Like there's still a lot of people haven't made the full migration across of their data cloud environment or figured out how to expose that to agents. Um but as they do that, you know, uh that's that's that's obviously going to become a a problem. The problem now is the agents probably don't have access to enough context to be able to give meaningful in you know six to 12 months time for a lot of organizations and there's already some in this spot.
[00:52:50]
Um they're not going to uh they're going to be having so much context the agent basically becomes useless because it's getting too many hits or they're filling the context window or they can't use cheaper models which typically have a smaller. Um, so I think being able to thin slice that and again I think graph uh helps with this because it u rather than you know rag or uh you know or uh just randomly stuffing in every bit of context into a system. Um the knowledge graphs allow you to uh to thin slice uh the information that you're putting in um through you know more more refined and target queries. So I think knowledge graphs are going to become more and more important for organizations or uh ways of smart indexing. You know I know Snowflake has has some of these solutions. data bricks as well um has have these and so you know uh ways of um of fast discovery of just the right information just in time and injecting that in is going to become really important and as you mentioned a case is going to become uh more and more of a you know an important concept for organizations which perhaps haven't had to think about it before because that's a cheap and fast retrieval uh and uh you know so I I think there's a you know a lot of space for you know the the the um companies you know are thinking horizontal scale uh and uh you know build it wider build it deeper um whereas I I think you know the cost is going to come just prohibitive and so they're going to have to start to think build it smarter um we're in that early stage where throwing money at a problem solves it eventually that starts a threshold and you and you're going to have to start throwing intelligence at it right and smart problem solve. Um so yeah I I think we're getting some companies are already there others are starting to get there and more and more you know we're in this sort of um you know this uh pyramid phase where got a handful of companies that are actually ready for that but 12 months time is going to be every consultants are going to be off their feet solving these problems. Yeah, hopefully. Uh, look, uh, as we start wrapping up, Adam, I do want to close on, um, the the the governance aspect of AI. Uh, you alluded to that already. Um, u, and the the level of governance varies on on AI when when you look at more regulatory industries versus more technology digital businesses. Um, I also like you, I also come from health insurance background.
[00:55:29]
So, I know the rigor of governance that is required in those industries. Uh and but the early stages of air governance was very much modeled on on on typical traditional governance models where you'll have a governance forum where the new AI opportunities will be discussed, the ethical aspects will be discussed um etc. um but I think um u I I'm sure you will agree that we need to now start think of the broader um uh definition of governance for AI. uh one as one it's all about not only the ethical aspect but the cost observity aspects etc the phops aspect um also what needs to be more real time versus what would be reactive right so how do you see the overall discipline of governance emerging or should emerge in the future >> yeah look um one one of the things I I really I really enjoyed it's a funny thing to say I enjoyed governance but um Uh I uh you know I'm I'm a move fast and break things sort of person. You know I like to I like to see if I you know I like to ask can I rather than should I uh and that doesn't really go that well in a you know in a heavily regulated um insurance space. Um and so well before I got into the role of um of AI architect they already started to put in the governance stream. Uh and so that really gave me a an operating model and framework to to to play within. Um which was fantastic and and you know it sort of predefined uh a lot of the structures that I had to build. uh and I think a lot of companies their their appetite is going to be to jump straight into building solutions uh rather than um first of all wrestling with some of the ethical and some of the um you know and some of the regulatory things that they they really need to cover off first. Um and what what they'll find is uh that'll become just a massive break like a a massive break for them unless they sort that out first. Um because you know uh we're playing with people's we're playing with people's data. We're playing anywhere in the health insurance space like you lose someone's you know uh license details or credit card information. You can issue them a new license. You can issue them a new credit card. You can't issue them a new health history.
[00:57:57]
>> And so uh being ensuring that there's security around um and this is something that that NIB was fantastic at. their cyber team was was just very good, very diligent and um and you know it it gave me some assurance as I was building um that that that was well governed data. Uh and if you really want to give and this was a real hesitation within engineers and and adoption in general. if there's if they don't have a confidence that the uh that the end systems are going to prevent them from making horrible mistakes, they're not going to adopt it because engineers are smart, right? They're some of the smartest people in organizations and um and they also want to do good, you know, they want to build good things, they want to build good products, they don't want to uh hurt the end customer. And so if there's any sniff of that, um they'll they'll be be resistant to adopt. And so companies who want really want to see AI adoption and their engineers pick it up and you know prevailing in their teams and you know building these amazing workloads and agents in production um they really have to think about governance as as the as the a priority first thing they jump into or their a or their engineers are just going to be so anxious around you know the potential negative externalities of what they're building that they just won't do it. Um and that will be the common you know the common thread of objection will be I can you guarantee me that this you know I I'll build this and ship it c can you guarantee me that you know that this will protect the end user um through the guard rails that we built on on the you know on the um on on the surrounds of this deployment and you have to be able to answer yes to that otherwise the engineers just they won't uh and so some of those guard rails need to be shifted left you know I spoke about governor module where uh you're injecting certain governance pieces at invocation time um asserting them on CI but absolutely on the back end of the governance making sure you uh you know every run of every agent what it did what its output was was that compared against the gold standard um you know was every action logged was every tool call logged uh is someone reviewing that or aically reviewing that at least like these are you know shipping something to production.
[01:00:23]
Building a prototype, you know, it can be 90% good. Um, and uh and that's good enough. You you have multiple agents, you know, typically in any system. Um, now that that compounds and that that compounding debt is now, but by the time you get production now, it's only 59% efficient. Uh, it needs one 90%. And so wouldn't you ship something that you were, you know, 59% sure wouldn't lead data? No. Of course you wouldn't, right? We we need better. We need to be thinking as organizations around these these guard rails, not to get everything to 100%. But have we built a system overall and this is for humans or for agents um that can that can give some assertion and some guarantee that we're not going to do wrong by our customers. Uh and I think that's um I mean that's an that's way more important to think of uh if you if an organization was thinking of um adopting AI both for the human engineer and the and the agent engineers um before they before they even turn on a claude license um you know and does does our system inherently protect um the outcomes of our customers and if and if they can't answer absolutely then you know they they stop shipping code now and fix that >> a lot of shipping agent.
[01:01:48]
>> Yeah. On that note, Adam, thanks thanks for your time. I think this was really wonderful conversation. Um, we covered lot many topics under AI engineering. Um, and people who are really interested who are thinking the in in the middle of uh um building the AI engineering practices or even in general think of how to embed AI into their their organizations. I think there are lot many takeaways for them. So thanks again. Thanks for your time. Pleasure hosting you. If you found this discussion valuable, please follow and subscribe to Enterprise Tech Talk. And thanks for listening. I look forward to seeing you in the next
This is where harness engineering becomes critical. Adam describes harnesses as the control layers that help AI tools produce enterprise-aligned outcomes. These may include policy-as-code, security standards, QA expectations, cloud patterns, coding conventions, architecture guidance, compliance rules and organisational context. In practical terms, the AI model is only one component. The surrounding harness determines whether the output is safe, reliable and fit for enterprise use.
The episode also explores the shift from vibe coding to spec-driven development. AI can quickly generate clickable demos and functional-looking prototypes, but enterprise software requires far more rigour. Production systems need clear product requirements, architecture thinking, solution design, test criteria, phased delivery and review processes. Without these disciplines, AI-generated code can accelerate technical debt rather than reduce it.
Another important topic is loop engineering. As AI agents move from single prompts to multi-step, goal-driven execution, teams must design loops carefully. Poorly designed loops can increase token costs, create context drift and generate large bodies of code that engineers struggle to understand. Better-designed loops use clear goals, sub-agents, model routing, validation steps and human review to improve both speed and quality.
Testing and quality assurance are also changing. In traditional software delivery, tests are often binary: pass or fail. AI-driven systems are more nuanced. Adam argues that QA must shift left and become more engineering-oriented, with stronger involvement in specifications, harness design, golden datasets, runtime evaluation, automated testing and end-to-end validation.
The conversation also highlights the growing importance of context management and enterprise memory. RAG and long-context models are useful, but they do not always capture the “why” behind technology decisions. Adam points to the need for graph-backed memory, decision ledgers and code context that preserve intent, ownership, change history and architectural rationale over time.
The episode closes with a strong message on AI governance. In regulated industries, governance cannot be a late-stage review or an ethics forum alone. It needs to be embedded into the AI engineering lifecycle through runtime controls, logging, observability, FinOps, policy injection and guardrails that give engineers confidence to adopt AI safely.
The core takeaway is clear: enterprise AI success will not come from model choice alone. It will come from the harnesses, guardrails, context layers, memory systems and governance practices organisations build around AI.
AI engineering is where enterprise AI moves beyond the hype and becomes production reality.
Episode Transcript
FULL TRANSCRIPT
This transcript is based on the episode’s English auto-captions and has been formatted for readability. Please allow for occasional transcription errors in names, acronyms and specialised terms.
[00:00:00]
Atlassian study engineers only spend 16% of their day doing um actual coding right and so uh we're really optimized at the moment for that 16% where we're filling with characters um and and that was never the job of engineering. I think a lot of organizations are stuck in the conversation around which model they're using. Um, which is a very, you know, 2025 discussion. Um, 2026 discussion needs to be around, you know, how are we empowering people with better harnesses and empowering the whole world with the meta harness that we're building. Just like you, you know, a human needs to uh needs to have um some framework around what it's building. This is what specdriven development obviously uh you know gives it and so um you know some people see specri development as you know just the PRD where it's you know a definition of of uh various pieces and I find that a little bit loose now nowadays QA is I think the the biggest sort of unsolved problem at the moment in uh in the SDLC agentic SDLC uh and I think it I think it falls it's a problem in two ways one is we've always talked about QA shifting left.
[00:01:15]
Um, and uh, the faster we move with with agent development, the less likely we are to actually shift the QA left. A lot of companies, their their appetite is going to be to jump straight into building solutions uh, rather than um, first of all wrestling with some of the ethical and some of the um, you know, and some of the regulatory things that they they really need to cover off first. Um, and what what they'll find is uh that'll become just a massive break, like a a massive uh break for them, unless they sort that out first. Hello and welcome to the Enterprise Tech Talk podcast. I am your host Somitra Kalikar. Now AI has moved rapidly from just being a curiosity to an enterprise priority. Many organizations are now actively exploring uh assisted software disco delivery and uh agentic workflows.
[00:02:26]
uh organizations are now discovering that um the building a flashy AI prototypes is vastly different than engineering secure, reliable and cost effective AI systems and that is why today's conversation focuses on AI engineering. It's not only about prom engineering or selection of new models but the broader discipline about how to build reliable AI systems in production. And to help me unpack this topic, I'm joined today by Adam Vidonoski. Adam has a wealth of experience at the intersection of AI software delivery and uh enterprise technologies and with his insights. Today we will explore how to move from a prototype hype to a production reality. Adam, welcome to the podcast. Great to have you here. >> Yeah, great to be here. Thanks for having me. Um, Adam, before we get into details, uh, tell us a little bit about your professional background and how you got into AI engineering in the first place.
[00:03:31]
>> Sure. Um, well, uh, I've been a been in software for about 30 years. So, I've been a builder. Uh, I've run my own companies, done startups, and and predominantly around that um sort of startup and and, uh, consulting kind of space. Um, NIB uh is where I was at last as a uh AI architect and prior to that uh I was in the same organization as a senior engineering manager um working with a you know a relatively large team um in the uh in the travel insurance domain moved across into the group domain for AI architect. uh obviously when AI started you know it sort of entered entered the uh the common man's world through chat GPT you know four four five years ago uh this um little chat box that we that uh it just it just seems like such a you know simple and almost pathetic way to interface with a profound technology and yet this is what we've landed on um is uh and sort of entered into the world and uh and immediate mediately I think like everyone um was just blown away by the potential of it. And so of course uh I started you know digging into it and and trying to understand how this new attention is everything idea works um looking at how you know the the underlying models work and and vectorization etc. and um and became quite enamored with that as an idea and started to sort of steer myself more and more towards that. As a senior manager, you tend to not be on the tools that much and I found myself on the tools every night. Uh you know, >> playing with it, understanding with it, using it for development and as APIs and whatnot became available, um building products with it. And so uh yeah, I've been playing with it for um you know, four or five years. Obviously following it along, um various other iterations of uh of ML prior to that. Um but particularly in my last role for the last 12 months, um running functionally a lab within uh within NIB doing experimentation and ultimately product delivery uh with uh with AI. Um so it's been a it's been like everyone a wild ride. changes every five minutes. Uh you have to be reading the news every two to keep up. And um it's it's it's been a lot of fun obviously.
[00:06:03]
>> Yeah. And as you said uh we all got exposure to LM based AI around four years back. Um uh but in the last one year or so I would say the the the AI engineering the uh that discipline has started becoming much more mature. Right. Um and do you see do you think we are at a stage in terms of maturity of uh AI based engineering and coding that it is now becoming its own distinct discipline compared to traditional way of engineering? >> Um yes and no. So the the the concepts are the same and so like I think if you were in an organization uh you know humans make mistakes all the time and uh and we make coding mistakes all the time and and we're you know perfectly for people all of us. >> Yeah. Um and uh if you're an assistant, if you're in a organization that has built good guard rails around uh around human engineers, then you're probably already directionally correct towards having uh you know an organization that's reasonably well or at least directionally set up for having um agents um also write code because you know good good environments are uh are environments that are intentionally low trust. So you know they we have a good PR discipline good CI/CD uh pipelines um ways of doing um you know validations on data that's that's coming in and data that's going out all of those are still required disciplines uh for an organization to have with AI u typically and this was a this was a Atlassian study engineers only spend 16% of their day doing um actual coding right So, uh, we're really optimized at the moment for that 16% where we're filling with characters. Um, and and that was never the job of engineering. The job of engineering has always been problem solving, architecting, um, and imagining new ways of solving problems. And so, I I think that that as a discipline is still very much at play. uh and uh and the you know optimizing around that 16% and the various other things that you need to build around that to make that um safe and and secure for deploying into an environment. um you know functionally ex exist but they need to be scaled very differently because you know if you can imagine if you've got 16% of an engineers's time you've got an organization with you know 100 a thousand engineers and um and all of a sudden that's through that same pipe you're now shoving 10x of the the code through obviously your PRs are going to break your QA processes are going to break uh and your um you're shoving things into that pipeline through um product managers, product owners, BAS, uh and their back end as well, your cyber security uh and your cloud.
[00:09:09]
They're going to start to strain under that. So, you do need to think about how do we widen the whole pipe, but um but functionally, I think we've we've thought through these problems um historically. We just need to figure out how to scale them in the same in the same way we've scaled uh the code generation. Yeah, I think uh it's wise to say we are building on top of all the fundamentals that we have built as part of the standard software engineering practices. Uh but that 16% that you said that is unique potentially uh that definitely I want to unpack in terms of what it means when it comes to the designing the reliable and predictable AI solutions etc. But before we go there, this space is definitely is filled with many hypes, right? So I just wanted to pick up your thoughts on at a high level which of the trends you see are um overhyped today in the AI engineering space and and more importantly which trends are have emerged but are underrated which are not talked about that much but should should have get that much more attention.
[00:10:19]
Yeah. Um, you know, you go on LinkedIn and, uh, every every man and his dog is on there saying, "I built Spotify over the weekend." >> And, um, that is, uh, good for you, brother. Guarantee you didn't build Spotify over the weekend. Uh, you built something that was perhaps a nice clickable UI. uh and if you try to launch that into production um you are going to you know you're going to lose your customers data data money etc. Um so they took pock and um and AI is incredibly good at building pocks. Like just this morning, I I I've had an idea for uh for a while that I've been going back and forth on. It's quite a complex um problem space to solve in and I thought I'll just I'll just start, you know, work start from the UI and work back towards the back end that I've started to build out. Uh, and I got a I got stunning incredible demo, you know, out of a PRD uh that that I've spent, you know, months um re refining, but um, you know, it it is it's not a product, you know, it's so far from a product.
[00:11:27]
Um, so there's definitely this this hype um because we're and it's it's addictive honestly like this, you know, how fast you can spin something out that is that is functional, clickable uh and uh and can solve a real problem. And in some cases that's enough. You know, if if you're solving a if you're solving a problem just for yourself, um you know, sometimes that is that that's completely sufficient. Um but if if you're going into uh a production environment uh you need you need far more infrastructure around that process uh than uh than you know I built Spotify over the weekend. Um and you know the other I think the other hype piece is that you know it's going to replace engineers in the next 3 months six months. I I don't believe that for for a moment. I I think eventually over time perhaps within the next two years it will I mean it's already reshaping um organizations and how they think about uh how they think about engineering and other disciplines within the organization um and over time it will dramatically change like without a doubt the the the um the the discipline um but not today probably and enterprises are typically slower to adopt and so you know that that transition take a while. Um but the the underhyped things I think are around the importance of um of harness uh and harness engineering uh and um and something which people aren't talking about as much uh two things. one is multiplayer uh aspects of of AI and your AI use cases um where we you know at the moment we've so we've spent the last couple of years training agents how to work as a team um how to uh how to swarm over uh over problems um and we've we've left humans out of that as this as this single player uh you know in through a through a chat interface um whereas I think we need to broaden that out into a much more multiplayer mindset where we incorporate other parts of the organization, other parts of engineering to be able to influence the outcome. Um, and people tend to talk about it as harness. I I talk about it as I met a harness or a harness harness because it's the things that outside of you as a personal developer working on something to start to include the whole team um to be able to shape just the influence that's happening on your machine. So um that's where I think we just need to start moving towards a conversation of multiplayer and and and uh and meta harness or harness harness. Uh but certainly at at an individual level um that you know that the harness is is an incredibly important idea. I think a lot of organizations are stuck in the conversation around which model they're using. um which is a very you know 2025 discussion uh 2026 discussion needs to be around you know how are we empowering people with better harnesses and empowering the whole or all with the meta harness that we built.
[00:14:42]
>> Exactly. And that was I wanted to unpack that anyway. So good you actually given the direction on that. Um um the if you had to take a deep dive the first point I wanted to unpack with you was about how we ensure reliability of u uh the AISS systems uh in production and uh as you rightly said um selecting model is just one fraction one one part of it um but that's not everything the more important when it comes to reliability aspects in particular is how you ensure the security how do you ensure the the the the overall reliability of of the system and building that harness around that becomes incredibly important and industry has shifted from the product the the prompt engineering to now focusing more on harness engineering. Right. So if would you mind unpacking that when when you talked about harness what really goes underneath what are the key control planes uh organization should be thinking of to build a good harness. >> Yeah. and and you know uh for when people talk about harness they're talking about multiple things and so it's just a it's a it's a catch all bucket at the moment and I think we need to start to distinguish um what we mean when we say harness because the harness could be um you know the the various tools and uh and memory types and things that you use in an agent >> um it can be you know clawed code and uh and uh codecs uh and um and cursor are all examples of harnesses in their own right. Uh but then when engineers start to talk about harness, they talk about the custom tooling they they build the sub aents um they get called during invocation. Uh and so there's that uh you know there's those ideas and then I've also seen people talking about harness in terms of the whole team um infrastructure that sits around uh you know the single individual engineers. Um when I when I talk about uh you know harness or I I typically talk about a meta harness because I like to think in terms of whole teams um an individual is going to build the way they build right and if they some of vs code or jet brains historically you know they pick their IDE um they you know engineers are going to build their their own coding harness uh setup and that's going to become their their secret source their magic you know like like a Jedi building a lightsaber uh and and it's going to be unique to them. But the magic wasn't ever, you know, if you look at Star Wars as an idea, uh you know, we we're enamored with the Jedi and their lightsabers. Uh and uh but that's just a that's just a tube with a with a crystal in it and some force magic. you know there is probably a bad example but you know the Death Star now that is a >> that is an architectural marvel you know and it's a what a stunning picture of civilization level uh collaboration and and so I think really the power going forward is going to be at that's you know at that sort of meta scale uh where we where we start to think about harness harnesses or meta harnesses um so you know some of the things you you mentioned cyber security um there's obviously things like QA say uh there's um there's cloud infrastructure marketing has a say legal risk compliance they all have a say uh in how things get built um and and uh models and coding agents are really good I mean really good right now you know since December and you you had um what was it Opus 4.7 uh come out around December uh they they've been they've been astronomically good at generating Um and and since you know you've had uh various GVT models come out and they're also exceptionally good. Now you've got Chinese models GLM 5.2 and uh and Kim K 2.6 all very good at generating uh code and so like personally I haven't I've barely written a lot of code in 12 months you know like I've written code of course but um it's it's probably down on that 10%.
[00:19:00]
Um, and when you're doing green fields, that's totally fine. Like if I'm if I just put in a put in a PRD or a prompt or something and have it generate some code, uh, and I don't care that much around, um, you know, style or language or, uh, infrastructure or databases, it's going to turn out something that feels like black magic. And that's, you know, that's >> every man and his dog on on LinkedIn who built Spotify. Uh, they're able to do that. uh but in terms of generating code particular to your organization uh that's a that's a completely different story. They don't know the the model doesn't know anything about your organization. And so this is where we we need to shape a meta harness around it to shape the outcome together uh that that's unique to the to the needs of your organization. And so you know one one thing that I uh built uh was it was called the governor module. Um and that allowed actually two things that I'll I'll I'll reference. One was governor module. Uh and that allowed um vertical teams. So uh teams like cloud cider um you know data every every engineer always screws data over because we you know we make changes to the model that we don't tell them. So they're adjusting you know broken pipelines etc.
[00:20:18]
um risk, compliance, legal, marketing, they can all define their policies uh and those get converted to policies of the code uh and injected um at invocation time into the coding agent either via hook or MCP. Um similarly, horizontal teams that are doing the actual can define their uh their um particular teams um style guides etc. and that gets injected at invocation time. And so now you have a control plane that sits over over everyone's coding agents that um that that's uh that that is getting injected at invocation time um the various policies of various parts of the organization and and this is what I call giving agents good steer um the engineer can do good steer for sure uh but when time crunch happens they they're probably not thinking um about the you know 150 page documents that uh is sitting there in confluence that define our cyber security mandates and and how we and the agent definitely is thinking about that.
[00:21:26]
Um but if we define policy as code and injected an invocation all of a sudden that shapes what gets delivered to be more organizational shaped and that can you know include things like how we deploy how you know what cloud infrastructure we prefer how we like to secure APIs etc. then um and then since you have that policy as code, it's very easy to then reassert it at CI and so uh you can you can virtually guarantee that um that what you know uh it needs to be the shape of code for your organization ends up getting delivered in a fairly similar shape to what you need for your organization because you've uh you've worked collaborative collaboratively as a team and if that changes over time or each of those teams can own that vertical injection. The other piece is um coding agents are very good at working on the individual codebase that you are reading. Um and so if you open up, you know, a GitHub repo and it it it looks at it and it's able to reason over it and that's great. The problem is organizations very rarely have a single codebase. Um NIB had 3,600 plus repos.
[00:22:37]
uh and you make a change in one of those repos, there's a high potential that there's going to be an upstream or downstream impact uh you know to at least data or another consuming service. Um and so we need a way to give visibility to the coding agent for uh repos that sit outside the the the current thing you're working on which you may not not even know as an engineer you're making downstream. Um and so uh something that I built uh which I called Skynet rather ominously ingested all the repos into a knowledge graph uh created a vectorization uh and um and extracted code intelligence um uh code intelligence database out of that as well. And so gave the agents uh the coding agents of the engineer a 360 degree view of all the other repos uh that it could pre-calculate um downstream impact and upstream consumption um before it made a change. And so, you know, neither of those are are cheap um things to build and neither and and there are some examples of these like source graph um has something like that for the for the Skynet.
[00:23:47]
Um but organizations need to start thinking about how they shape that meta harness to create good steer for the coding to make organizational shaped code, not just code in general. Yeah, clearly you are a huge fan of Star Wars and Terminator. So, uh we we are definitely aligned there. Um look uh and I think uh one thing you said uh and it's I mean if I had to summarize is traditional soft in the traditional software engineering most of those control planes in a in a way where more reactive in the sense that you build a code then someone will go through a a security governance review. you'll go through any of the different types of reviews but that is typically much more reactive um uh u in as in the AI world when you are talking about the 10x engineer uh uh workload um and uh the the work is getting accelerated you need to shift from reactive to more runtime controls right and then that's where that harness becomes important that it has to be more runtime um than being reactive uh But my question also is given given this is the emerging capability need uh do you see industry solutions emerging as best of breed those runtime harnesses um or do you see organizations still need to build their own harness models?
[00:25:18]
>> Yeah, it's um it's fragmented. Uh so you know going back 12 months when I started the AI architect role there just wasn't a lot of these solutions. You sat to you had to roll your own and they fit into a category of what I call boring problems. Um and so you know uh doing doing um PRs is a boring problem. doing agentic PRs now is becoming increasingly a solved problem because it's it's the clear next bottleneck after shipping right is you're you're gonna uh you're going to ship a large volume of code and it needs to be PR and historically uh those have been poor anyway right you know we we tend to in engineering rubber stamp uh stuff you know looks good to me uh approve um which is a bad way to do PRs but it's probably how a lot of PRs happen in companies anyway. Um, you know, and so the the agentic PRs uh and there's some debate about how best to do this and uh you know, should it be the same model as doing the invocation and you know, whatever you we can debate that um till the cows come home, but uh at the velocity that we're that we're going to see code getting shipped um that's 100% need to is going to need to be authentic. Um, and so in some s slices of of the uh of the um SDLC, we're starting to see these tools uh emerge. Um, and I and I think that's, you know, that's great and and and but companies are already starting to cry foul at the cost of tokens.
[00:26:53]
>> Um, these are these are just more token consuming machines, right? Right. And so if we're not comfortable with the invocation tokens, we're probably not going to be comfortable with PR tokens, with the cyber tokens, with the QA tokens. Um, and so I think we need to, um, we need to either get comfortable with, you know, with token spend or figure out better ways to squash the cost of tokens or get more comfortable using, you know, open source models, uh, and model variable model um, uh, variable capable models. So, you know, throwing out some bits to say Hi coup or Gemini um 3.5 Flash for cheaper operations uh rather than throw everything at Fable or or >> Okay. >> Uh and that's typically, you know, not not handled that well in organizations. Everyone's throwing stuff out to superm models as opposed to, you know, figuring out these nice thin slices.
[00:27:51]
>> Yeah. Let's let's shift our focus on um uh on the SDLC the software delivery life cycle approaches emerging in in AI engineering and and when we started AI based coding as you know um we started with just asking uh AI models chd or sorry codex or um uh claw to build a small feature build a small tute a small task and what anthropopathy that time coined pipe coding and um uh we and that's the uh term that has stuck with us. Um but as we have mature the AI engineering practice um the shift has uh uh has been around uh not just giving a task but doing a proper uh spec based development specular development so the so that um uh the AI doesn't do the coding just by intuition but it has a very clear view on what is the requirements what is the inputs and how it should test the outputs etc. And now we are also showing the shift towards more goal-driven development. Um right which is you don't necessarily go and intervene at every stage but you define design a very clear goal and and then let AI iterate towards achieving that goal. How do you see this these shifts emerging and particularly I want to get your thoughts on when we talk about specdriven development what best practices are evolved there and how spec development and goal-driven development which is now being taught more nowadays how they're going to complement with each other >> yeah sure so I I I before it had a name I was doing spectrum development I my first my first use of AI for coding was copy and paste like a lot of people right you know you ask a question and you know basically using using chatb like stack overflow or something is spitting out an you copy and paste that into your IDE and run um you know obviously that was inefficient and so uh companies like cursor and uh and claude code uh saw that you know single line complete and and kudos to GitHub copilot you know they were very early on this uh and they were doing this single line completion um and uh And you know that was great uh for for what it was. Then the others thought what if it could just complete more than one line? What if it could complete a whole file? What if it could reason over multiple files? And so cursor and cloud code come in and start to do that uh and say do the copy and paste as well larger chunks. Um obviously as you start to handle more or hand off more to the coding agent uh it needs more context. And so uh you know if if you ask it uh you know build me Spotify um it has a it has a really great reference to go from it can it understands what Spotify is. It understands the branding. You can probably go off and search it and find you know the uh the DOM and uh and return you something that's pretty similar. Uh in most cases you're not going to be rebuilding Spotify. you're going to be building something particular to your organization, uh, a feature that's particular to, uh, some functionality that exists in your software. And so its references are going to be, you know, more limited. And so giving it just like you, you know, a human needs to, uh, needs to have, um, some framework around what it's building. This is what specdriven development obviously uh, you know, gives it. And so um you know some people see spectrum development as just the PRD whereas you know a definition of of various pieces and I find that a little bit loose now nowadays um whereas I like to use um in that context I like to use uh because I also like testing um I've been experimenting more and more with say girken and and uh and cucumber style practices where you give it uh you know smaller smaller chunks which become ends up becoming tasks Um, but I also like to then go through and do an architectural piece and a uh and a solution design piece. And so I I can get more granular in terms of the steering, the upfront steering of what's going to go into the coding agent. Now um and so but I would consider spec uh to be you know really clearly defined um uh you know thing that's going to be built. So a product requirements document um with some breakdown of um of the of the end state uh user you know usage requirements and and what basically good product practice uh and then um and incorporate some early injection of u of QA practice using uh you know things like cucumber and girkin style language um and then move into uh the architecture and so I'll have a separate architectural document um where I define um you know it might be uh I get the original architecture of a piece and then extend that uh and you know that gives me some assurance that it's going to land near where I want it to land uh and then a solution design dock similar to what I do and and honestly the models are very good at generating almost all of that as well. So your spec >> generate lots of it, review it, and at least you have reviewed.
[00:33:14]
Um, anyone who's been doing specri development for a while knows that the practice ends up becoming, okay, I give this to Claude. Uh, Claude, you know, does and I always do, you know, a phase delivery. So give me, you know, I want this to be a phase delivery, multiple tasks per phase. Um and uh you know and and it will split out that those phases into tasks and then quite happily tick through those tasks. Um you know what happened about six or seven months ago people started talking about loops uh and you know this obviously the Ralph Wigum loop and others um uh things like uh Gas Town um which automate you know large portions of this process. GSD is another one. Um and uh and so the what you what they found was or what engineers were finding was that the model would go through or their cord would go through as an example um and build task one, task two, task three and it's like are you ready for task four?
[00:34:13]
And so their their next action was yes enter you know and so their involvement in the development process and ended up becoming saying yes and hitting enter. uh and uh and hopefully they had some other you know um skills and things in there or some uh something in their claw MD or what have you. Uh they were steering the agent to also build tests and other things. Um and so you know their their evolvement became less and less the better the models became. And so then you start to go well why bother hitting yes enter at all? Why wouldn't we just loop through that um and get you know through full completion? Um now what I found with what I found with a lot of the sort of uh you know the particularly the early loop test and now there's now there's um you know some more defined golf um earlier models would hit about a 70% good outcome. Um, now we're getting to more like an a probably an 80 to 85% good outcome. And you still need to do a lot of, you know, uh, rework and testing and and but most of it I find now at the end of the loop or or at the end of a goal is I'm educating myself on what was built and getting it to replay to me and prove to me that what I've built was good. Um, but these, you know, I've had loops run for for just, you know, hours, whole days. uh and at the end of it a I'm poorer uh because the token costs uh but b um have a just a huge corpus of code that now I have to try and understand um but functionally uh more and more and particularly with you know models like fable and and gtd 5.6 six. Um I'm seeing that there's there's code that is that that is just very good or or outcomes that are very >> and and the sorry no but uh so the the loop engineering is definitely becoming uh attractive and uh there's a lot of optic to that but there are also a couple of concerns and and one you just mentioned is um if not properly designed it can lead to token maximization. Um that's one thing and second thing is uh as you keep uh as AI keeps running into loops and the context keeps piling up building up um uh the more context that is built there is always um a risk of drift u and um how so so what are the best practices that can are emerging or what are the guard rates should be put in place in designing the loop so that these concerns are are addressed.
[00:36:52]
>> Yeah. So um one just I mean this is a really simple one a lot of people already doing it is um is you know first of all sub agents obviously in your loops and how you define those and um and so you know sub agents will obviously get their own context and so uh so heavily heavily relying on sub agents um but then having a I like to use a good model as like the orchestrator model that sits over the top something you know now fable obviously is a good choice is to define all of the tasks and define all of the um uh you know the the what's ultimately going to get built. Um hands off to smaller models like Sonnet. uh Sonnet 5 is very capable um at generating code particularly if it's going to be validated by a larger model like Opus or or Sonnet um and uh and in some cases you can have um even Haiku uh do do things like documentation and whatnot which are also uh by volume expensive exercises right uh and uh um and so writing tests or running tests or u or running the code and watching the output haiku's great for that. Um, and so I think you know the the cost savings I mean there are a couple of a couple of ways. One is um in organizations this is a little bit harder because your cash pool is often smaller if you're running out of like an enterprise bedrock environment um but but trying to optimize for c hits and and and that sort of thing. squashing um squashing your uh your inference uh inputs and outputs with something like tune or JSON reduction in some way. Uh you know we've seen 30% plus reductions um there. Um and so you can minimize the the raw token burn, but then uh and the the raw query burn um cost you can you can definitely do by by using um sub agents that you clearly define um and uh and hand off a lot of tasks to smaller cheaper models. Um but I typically like to have a larger model driving um and then in post validation uh doing the doing the code review.
[00:39:06]
>> Yeah. Look um the other other uh uh point I wanted to unpack with you was around u uh the the testing uh and the evals for for testing. Um now in in the traditional uh software engineering and testing the traditional safety systems is all about you have clearly defined test cases test criteria you run those test cases after feature is built and um and you you either pass or fail essentially the test case right that's not as binary um in in the AI AI AI driven AI systems right um how what what are your um uh recommendation in terms of building good test practice to support um AI based developments. What are the runtime events which um engineers should be thinking of? Yeah, I think you know out of the full SDLC you know uh cyber used to be the one that worried me the most uh because obviously you know there's high risk there uh but but you know companies like whiz and and others have come in and they've built really great tool sets around that space and and there's continuing work happening there and and and of course um you know big companies have seen the risk and AWS is releasing stuff Microsoft Google and Um and and so that's less of a that's of an unsolved problem. QA is I think the the biggest sort of unsolved problem at the moment in uh in the SDLC agentic SDLC. Uh and I think it I think it falls it's a problem in two ways. One is we've always talked about QA shifting left. Um and uh the faster we move with with agentic development, the less likely we are to actually shift the QA left, right? uh we we actually because the engineer can just run ahead and and they often don't talk to the um to the QAs on the team um around how this is going to be tested you know what are the what are the uh you know what's being built and often the engineer doesn't even know that right they might have an idea from the spec but they don't know exactly the shape of it um until it's built um and particularly in an environment where you're uh giving less steer up front um so I think um you know in in terms of teams and enterprise teams, they they if there if if there's ever a time for shift left uh for for quality assurance, it's now. And so that needs to be quality assurance over the spec, it needs to be quality assurance over uh over the meta harness and understanding how that's going to deliver what it delivers. Um but more and more the the job of a QA is going to be less around um is less around testing and more thinking like an engineer and acting like an engineer. Uh and uh definitely at NIB we saw we're seeing the best QAS uh who were keeping pace with the um with the with the AI engineers were the ones that were thinking and acting bike engineers themselves and building um their own harnesses around uh around what's getting built. Obviously you got things like golden eval golden um golden data sets to eval against uh you know uh um building um of course most most of the time models won't volunteer to build u unit tests and u integration tests and um and so I think measuring that the test volume of that certainly um you know playright or or other endtoend uh testing um frameworks and practices are increasingly important and co Claude I'm sorry Anthropic has said multiple times um giving the agent the ability to test its own output is really important um and often engineers don't include that in their harness and I think that is that's a that's a real problem um definitely uh involving the QA very early uh and increasingly they're going to become engineers in of their own right building their own um or and this has always been true. There should be a tension between QA and engineers, right?
[00:43:21]
It almost um uh like the the QA their job should try and be to just fundamentally break anything the engineers build. Um and uh and I think their ability to do that they're I mean they're becoming just so outgunned uh with what you both the pace and and um and the visibility that that's more and more they really are going to have to be engineers who deeply understand the spec um and are able to build frameworks at the same pace as the engineers of building code. >> Yeah. Cool. Um uh the other um uh point I wanted to unpack with you Adam was around um the context management and the memory um enterprise memory management you uh and as you know you and I know uh that this is a unique thing about agentic systems that they are more effective um when the context is right u right uh that is not something that was uh relevant for traditional deterministics software. Now we um the the so the the industry has evolved on that as well in the sense that um we started with uh designing um the retrieval augmented model model applying retrieval augmented models to get the data organized context into uh the software those systems. Um but now we are also seeing um the emergence of large and availability of large context windows um for this uh agents right. Um where do you sit in in the in this debate uh whether rag is more effective or large context models are is is the way to go or there is something um they need to complement with each other.
[00:45:16]
Yeah, I think it's I think it's um actually a third option. Uh so, you know, rag rag is great if you want to uh if you want to understand does something exist, you know, in um in a document. Uh and um and but it doesn't give you why it's there. It doesn't give you um who you know made a decision. Um and you know there's of course things like you know I think PRD become if the repo becomes the source of truth PRD should be shipped with the repo uh and um and that becomes you know something like the why a change was made now managing that over time becomes complex right because you you in one one way of doing it I've seen people do is they just you know they update the PD to match the uh the um the code that's shipped and so you have a record in the git in the git history of changes over time um and why things exist where they exist. That's a you know and then the other side is people store feature um you know feature level PRDS uh and so they have a history of change over time um but the the context that that really matters and missing often from you know enterprises now that aren't using agentic delivery and it exists usually only in the head of a couple of senior engineers is why the codebase looks the way it does. uh why was change made um you know 14 months ago that's now causing us this problem in production I mean you have to go and ask this guy that's been there for 15 years uh the problem is these people leave and so that knowledge leaves with them and so I think there's an opportunity here with uh with you know agentic delivery to start to think about this problem again as a harmless problem or as a meta meta harness problem is capturing um the intent not just the you know uh the the change that we're going to make but the intent who who made it who made that decision why that decision was made um the point in time when that became a a a depre a deprecated idea um because you know like human memory we we learn new things and and we degrade old memories over time um I think uh so the third option I think is some combination of larger models for sure but you know I think that that's going to threshold um and you see that already like in a you know in a million token context window um you're really only getting you know sort of 500 um thousand tokens worth of useful uh memory there and after that it starts to degrade quite a lot in fact it starts to degrade you know after about a quarter the way through we start to see it you know slip slip a bit um and so I think we need to refine what gets passed in rag does that in part but as I said it doesn't give you any of the any of the the why or any of the when or any of the um you know the decision ledger or change in decision over time. So I think we actually need to have some combination of rag and graph um where the graph is uh is um the nodes are keeping track of um decisions over time um and the specific parts of uh of you know because you rag a you rag a a document or rag you know memories um and uh and you're going to return this tons of results uh and those may or may not be the context that you want. Um and so then you're again throwing that into the model to decide whereas a graph is going to be a little bit more deterministic around the choices that are made. And so you're starting to see companies like me zero and and others um enter into this space where they're uh where they're building um graphbacked uh you know memory systems and I think more and more this is going to become something which organizations want and need. Not only are they going to have their business context, but they're going to have this code context >> uh where you know we're going to be capturing the episodic memories out of every coding agent uh and and starting to store that in some way uh and referenceable through a graph so that every my coding agent has the same context as your coding agent the changes that were made over time. Uh obviously that's not really in place yet. You make a change in the code, it exists in your context. um and I don't have it and as soon as you start a new context um it's lost forever right and so we don't have any history of the decisions that were made we don't have any you know why that decision was made who made that decision and these are actually really important for organizations to have not just for you know for reasons for blame which we you know humanistically we like to do uh but for for reasons of um of context particularly as we get to you know uh which we haven't even spoken about yet which is you know regulatory and compliance problems particularly in the FSI space um you know where where is the human being in this process uh who drove these decisions why is the code shaped this way uh and so we're left being archaeologists of code without really having a map to be able to go through the history of the changes and so I think yeah graph graphback memory um where where we start to record sort of that episodic movement of of of invocation through through coding agents and uh agents running in the cloud etc.
[00:50:46]
um is going to become more and more important over time. Um so yeah I I think it's a combination of both but a third piece which I think is being able to actually point and time stamp it um with the decision and who made the decision. Yeah. And as organiz organizations are at the early stage of building designing and building these um context complaints and enterprise memory layers and and as you mentioned episodic is one of the layers etc. But I see as organizations start building more and more agentic systems um they will have to start considering on uh about how to optimize the context. For example, because >> um uh how to cache memory um because uh u different agentic workloads may require different context and just enough context that is also important because too much of context may also drift um the agent. So how do you see that those emerging concerns may not be the concerns today for organizations because that's so early stage but as this maturing these will become real concerns for them and how they should be thinking of addressing. Yeah, definitely. I already think I already think a lot of engineers just as an individual level over stuff um you know over stuff their coding agents and and I think there's always a danger that people want to over stuff their you know their cloud workflows as well like agents running better etc. Uh the reality is the reason it's not a problem for a lot of organizations right now is that they're that their data in general isn't ready. Like there's still a lot of people haven't made the full migration across of their data cloud environment or figured out how to expose that to agents. Um but as they do that, you know, uh that's that's that's obviously going to become a a problem. The problem now is the agents probably don't have access to enough context to be able to give meaningful in you know six to 12 months time for a lot of organizations and there's already some in this spot.
[00:52:50]
Um they're not going to uh they're going to be having so much context the agent basically becomes useless because it's getting too many hits or they're filling the context window or they can't use cheaper models which typically have a smaller. Um, so I think being able to thin slice that and again I think graph uh helps with this because it u rather than you know rag or uh you know or uh just randomly stuffing in every bit of context into a system. Um the knowledge graphs allow you to uh to thin slice uh the information that you're putting in um through you know more more refined and target queries. So I think knowledge graphs are going to become more and more important for organizations or uh ways of smart indexing. You know I know Snowflake has has some of these solutions. data bricks as well um has have these and so you know uh ways of um of fast discovery of just the right information just in time and injecting that in is going to become really important and as you mentioned a case is going to become uh more and more of a you know an important concept for organizations which perhaps haven't had to think about it before because that's a cheap and fast retrieval uh and uh you know so I I think there's a you know a lot of space for you know the the the um companies you know are thinking horizontal scale uh and uh you know build it wider build it deeper um whereas I I think you know the cost is going to come just prohibitive and so they're going to have to start to think build it smarter um we're in that early stage where throwing money at a problem solves it eventually that starts a threshold and you and you're going to have to start throwing intelligence at it right and smart problem solve. Um so yeah I I think we're getting some companies are already there others are starting to get there and more and more you know we're in this sort of um you know this uh pyramid phase where got a handful of companies that are actually ready for that but 12 months time is going to be every consultants are going to be off their feet solving these problems. Yeah, hopefully. Uh, look, uh, as we start wrapping up, Adam, I do want to close on, um, the the the governance aspect of AI. Uh, you alluded to that already. Um, u, and the the level of governance varies on on AI when when you look at more regulatory industries versus more technology digital businesses. Um, I also like you, I also come from health insurance background.
[00:55:29]
So, I know the rigor of governance that is required in those industries. Uh and but the early stages of air governance was very much modeled on on on typical traditional governance models where you'll have a governance forum where the new AI opportunities will be discussed, the ethical aspects will be discussed um etc. um but I think um u I I'm sure you will agree that we need to now start think of the broader um uh definition of governance for AI. uh one as one it's all about not only the ethical aspect but the cost observity aspects etc the phops aspect um also what needs to be more real time versus what would be reactive right so how do you see the overall discipline of governance emerging or should emerge in the future >> yeah look um one one of the things I I really I really enjoyed it's a funny thing to say I enjoyed governance but um Uh I uh you know I'm I'm a move fast and break things sort of person. You know I like to I like to see if I you know I like to ask can I rather than should I uh and that doesn't really go that well in a you know in a heavily regulated um insurance space. Um and so well before I got into the role of um of AI architect they already started to put in the governance stream. Uh and so that really gave me a an operating model and framework to to to play within. Um which was fantastic and and you know it sort of predefined uh a lot of the structures that I had to build. uh and I think a lot of companies their their appetite is going to be to jump straight into building solutions uh rather than um first of all wrestling with some of the ethical and some of the um you know and some of the regulatory things that they they really need to cover off first. Um and what what they'll find is uh that'll become just a massive break like a a massive break for them unless they sort that out first. Um because you know uh we're playing with people's we're playing with people's data. We're playing anywhere in the health insurance space like you lose someone's you know uh license details or credit card information. You can issue them a new license. You can issue them a new credit card. You can't issue them a new health history.
[00:57:57]
>> And so uh being ensuring that there's security around um and this is something that that NIB was fantastic at. their cyber team was was just very good, very diligent and um and you know it it gave me some assurance as I was building um that that that was well governed data. Uh and if you really want to give and this was a real hesitation within engineers and and adoption in general. if there's if they don't have a confidence that the uh that the end systems are going to prevent them from making horrible mistakes, they're not going to adopt it because engineers are smart, right? They're some of the smartest people in organizations and um and they also want to do good, you know, they want to build good things, they want to build good products, they don't want to uh hurt the end customer. And so if there's any sniff of that, um they'll they'll be be resistant to adopt. And so companies who want really want to see AI adoption and their engineers pick it up and you know prevailing in their teams and you know building these amazing workloads and agents in production um they really have to think about governance as as the as the a priority first thing they jump into or their a or their engineers are just going to be so anxious around you know the potential negative externalities of what they're building that they just won't do it. Um and that will be the common you know the common thread of objection will be I can you guarantee me that this you know I I'll build this and ship it c can you guarantee me that you know that this will protect the end user um through the guard rails that we built on on the you know on the um on on the surrounds of this deployment and you have to be able to answer yes to that otherwise the engineers just they won't uh and so some of those guard rails need to be shifted left you know I spoke about governor module where uh you're injecting certain governance pieces at invocation time um asserting them on CI but absolutely on the back end of the governance making sure you uh you know every run of every agent what it did what its output was was that compared against the gold standard um you know was every action logged was every tool call logged uh is someone reviewing that or aically reviewing that at least like these are you know shipping something to production.
[01:00:23]
Building a prototype, you know, it can be 90% good. Um, and uh and that's good enough. You you have multiple agents, you know, typically in any system. Um, now that that compounds and that that compounding debt is now, but by the time you get production now, it's only 59% efficient. Uh, it needs one 90%. And so wouldn't you ship something that you were, you know, 59% sure wouldn't lead data? No. Of course you wouldn't, right? We we need better. We need to be thinking as organizations around these these guard rails, not to get everything to 100%. But have we built a system overall and this is for humans or for agents um that can that can give some assertion and some guarantee that we're not going to do wrong by our customers. Uh and I think that's um I mean that's an that's way more important to think of uh if you if an organization was thinking of um adopting AI both for the human engineer and the and the agent engineers um before they before they even turn on a claude license um you know and does does our system inherently protect um the outcomes of our customers and if and if they can't answer absolutely then you know they they stop shipping code now and fix that >> a lot of shipping agent.
[01:01:48]
>> Yeah. On that note, Adam, thanks thanks for your time. I think this was really wonderful conversation. Um, we covered lot many topics under AI engineering. Um, and people who are really interested who are thinking the in in the middle of uh um building the AI engineering practices or even in general think of how to embed AI into their their organizations. I think there are lot many takeaways for them. So thanks again. Thanks for your time. Pleasure hosting you. If you found this discussion valuable, please follow and subscribe to Enterprise Tech Talk. And thanks for listening. I look forward to seeing you in the next
