開場:AI 如何進入 SRE
Welcome to the Podcast: AI in SRE
Ryan Donovan 介紹 Komodor 的 Asaf Savich,本集聚焦 AI SRE 所需的 context engineering。
Ryan Donovan 與 Komodor 的 AI Engineering Group Manager Asaf Savich 討論 SRE 如何運用 context engineering、sub-agents 與評估層處理跨服務事故,並讓人類工程師轉向策略與 AI agent 管理。
時間碼為章節起始位置;逐句稿不偽造 Spotify 未提供的精確時間碼。
Welcome to the Podcast: AI in SRE
Ryan Donovan 介紹 Komodor 的 Asaf Savich,本集聚焦 AI SRE 所需的 context engineering。
From Accounting to AI Engineering Management
Asaf 原想成為會計師,卻因 QA 工作愛上軟體,之後攻讀 computer science 並投入 startup。
The Growing Context Burden for SREs
Microservices、網路層與 cloud infrastructure 讓一次事故橫跨更多服務,SRE 必須同時理解技術、業務和組織背景。
Effective Context Engineering for AI Agents
可靠的 AI SRE 必須同時改善資料收集/注入與 evaluation layer;只增加 context 而沒有驗證並不足夠。
Delegating Tasks with AI Sub-Agents
主 agent 保留目標與決策脈絡,把 log summarization 等小任務交給專用 sub-agents,以降低 token、延遲和 hallucination。
The Sub-Agent Approach for AI Efficiency
即使大模型能容納數十萬 token,準確度仍可能下降;用較小 context 和針對任務選模,能更接近 production 所需的可靠度。
How to Guide and Trust LLM Judges
團隊從人工評估建立可信 judge,以固定 input 做大規模 A/B testing;成對比較比單獨打分更一致。
Building Training Incidents from Real-World Scenarios
客戶資料經匿名化後只保留問題本質,再用帶有大量 noise 的測試環境重建事故,形成每日 regression testing。
Autonomous Remediation with Human in the Loop
Claudia 能調查事故、提出 command 和證據,但先由使用者核准;低風險 rollback 比直接套 code fix 更容易獲得信任。
Identifying Fragile Infrastructure for Preventive SRE
AI 不只修復當下,也會回頭尋找深層 root cause;重複出現的 GitOps、database 與 networking 問題可交給專用 sub-agent。
Generating Post-Mortems and Learning from Incidents
Past postmortems 與 memory layer 能改善後續調查,Claudia 也可作為客戶 orchestrator 背後的專業 brain。
The Build vs. Buy Dilemma for AI Tools
自建或開源方案容易做出亮眼 demo,但在 noisy production incidents 裡,成熟的專業工具與支援仍有價值。
SRE's Evolving Role: Strategy and AI Management
AI 將主導繁重調查,人類 SRE 則轉向成本、規劃、review、approval 與 agent management。
Populist Badge Winner and Podcast Farewell
節目表揚 Joe Maller 的 Python which-command 回答,主持人與來賓留下聯絡方式並道別。
時間碼會跳到該句所屬章節;中文以傳意為主,技術詞保留英文。
These two must live together. One without the other is just not enough.
資料層與評估層必須共存,缺少任何一個都不夠。
Instead of having our main agent collect all the information, we like to delegate.
我們不讓主 agent 收集所有資訊,而是選擇分派工作。
Ninety-eight percent is just not good enough.
在 production incidents 裡,98% 的正確率仍然不夠。
Having good judges is the base for everything in the pyramid of velocity.
可靠的 judges 是整座開發速度金字塔的基礎。
We do all of that, but with human in the loop.
我們可以完成整個修復流程,但仍保留 human in the loop。
AI will lead the investigation and the SRE will be some kind of manager.
AI 會主導調查,而 SRE 將成為某種管理者。
事故可能橫跨數十個 microservices、logs、metrics、traces、程式碼、部署與商業背景,單一 SRE 還常需向其他角色取得資訊。
資料量增加會提高 token 成本與延遲,也可能降低模型準確度;production incident 對少量錯誤的容忍度非常低。
Main agent 保留事故目標與整體決策脈絡,專用 sub-agent 處理 log summarization 等小任務,再只回傳壓縮後的相關結果。
先以人工評估校準,保存固定 input 反覆比較,偏好直接問 A/B 誰較好,並用數千到數萬次樣本取得信心。
Rollback 風險較可控、可逆,比發布新版或直接改 code 更容易讓使用者建立信任;高風險動作仍保留人工核准。
AI 會承擔大量調查與資料整理,人類則更專注成本最佳化、長期規劃、策略、review、workflow approval 與 agent management。
英文以 Spotify 自動逐字稿為主,並以 Stack Overflow 官方頁面確認主持人、來賓、公司名稱與 ART19 音檔。已校正可確定的 SRE、Komodor、Claudia 等術語;標記片段仍可能需要對照原音。
Hello and welcome to the Stack Overflow Podcast, a place to talk all things software and technology. My name is Ryan Donovan, and today we're talking about context engineering in terms of AI SRE. My guest for that is Asaf Savitch, AI Engineering Group Manager at Komodor. So welcome to the show, Asaf.
歡迎收聽 Stack Overflow Podcast,這裡專門討論軟體與科技。我是 Ryan Donovan,今天要談 AI SRE 情境下的 context engineering。來賓是 Komodor 的 AI Engineering Group Manager Asaf Savich。Asaf,歡迎來到節目。
Hello. Hey, Ryan. Thanks for having me.
Ryan 你好,謝謝邀請。
Of course. So before we get into the context engineering discussion, tell us a little bit about yourself. How did you get into software and technology?
在進入 context engineering 前,先介紹一下你自己:你是怎麼進入軟體與科技領域的?
So I wanted actually to be an accountant. That's what I started off with. Yeah, I like numbers, I like money. I thought that was a good combination. And this is where I wanted to be. But then like I was looking for a job and like a QA opportunity came in my way and I started doing that and I just fell in love with software and just felt right and felt the, the natural thing for me. So I went to school and studied like computer science and it made me even more fall in love with, with the topic. And I just during my, my studies, I started working as a software engineer part time in a few small startups. And then I also understood the startups are also the right fit for me. So that was my beginnings.
我原本想當會計師,因為喜歡數字也喜歡錢,覺得很適合。後來找工作時遇到 QA 的機會,一做就愛上軟體,感覺非常自然。我去學校讀 computer science,求學期間也在幾家小型 startup 兼職做 software engineer,並發現 startup 同樣很適合我。這就是我的起點。
The software industry as a whole is sort of feeling the effects of the small decisions as they push vibe coded code to production. And I think a lot of folks are looking at ways to do the DevOps operations management.
整個軟體產業正在感受到許多小決策累積的後果,尤其 vibe-coded code 被推上 production 後,大家都在尋找更好的 DevOps 與營運管理方式。
So let's talk a little bit about using AI in SREs. We've talked to a few folks who, you know, do that sort of thing. Code in production is a big topic, right? Like what does SRE cover?
談談 AI 在 SRE 裡的用途。我們訪問過一些從事這類工作的人;production code 是很大的題目。SRE 的範圍究竟包含什麼?
So let's start with what's necessary actually is right Today in many organizations, SRE is a very pivotal position. SRE stands for site reliability engineer personally the guys who are responsible for keeping like the production environment stable, healthy and to make everyone happy, right? Both the the company providing the service and the users who are actually consuming the service. So SREs have that job and it's a lot of responsibility to do the job. And their part includes so much context to it. Even much before the AI era. It includes so much context to it to understand the business side of things, where those things fit. What's the relationship between services 1 to another. And over the years that we added a lot of capabilities of micro services and a lot of networking layer and cloud infrastructure. So we've just kept adding a lot and a lot on the shoulders of the SREs. We just have to make sure that production is actually reliable. So we put a lot of effort and a lot of responsibilities over the shoulder and just keeping adding those. It's just a bit problematic. So AI is a really good like to helping the SREs and making them be much better.
在許多組織裡,SRE 是關鍵職位。Site reliability engineers 要維持 production environment 穩定健康,讓服務供應方和使用者都滿意,責任非常重大。即使在 AI 以前,工作就需要大量 context:理解商業面、各項功能的位置、服務間關係。多年來又加入 microservices、networking layers 和 cloud infrastructure,不斷增加 SRE 肩上的負擔。他們仍必須確保 production 可靠,持續加責任並不可行;AI 可以幫助 SRE 做得更好。
That's a great point that you know, as people have sort of created micro services as a way to like organize teams like SREs have to know about all of that. That is a lot of information to have all at once, right?
Microservices 常被用來組織團隊,但 SRE 卻得了解所有服務。要同時掌握這麼多資訊,負擔真的很大。
With AI agents, one of the big topics now is context. How do you get context for all these AI decisions? And you know, we're just saying like there's a lot of context for SREs in general. How do you get that context into an AI and have it be effective?
AI agents 現在的一大焦點是 context:如何為每個 AI 決策取得所需背景?既然 SRE 本來就需要大量 context,要怎麼把它有效提供給 AI?
So we started off in a very like naive way of just, you know, throwing a lot of stuff they are and start to understand like what the agent is capable and is not capable of doing and really fast. We got to a point where it was just too problematic. And the data we had to inject into the AI agent, we had to think about it really hard and really long on like what's the proper way of injecting the data? What kind of information is more valuable than others? So we had to do a lot of stuff. So we focused on 2 main areas. One is actually the data and how to collect and to inject it. And the second is the evaluation layer that no matter what change do we think is important or seems the right way to go. We have an evaluation layer which each change that we are conducting, we want to make sure that the change is in fact putting us in the direction that we want to head. So these two must live together. 1 without the other is just not enough. And it's not enough for every AI agent, especially AI agents to that claim to be the ones that are actually resolving production issues and putting us in like the most sensitive place in the organization.
一開始我們很天真,把大量資料全丟進去,藉此了解 agent 能做與不能做什麼,很快就發現問題太多。我們必須仔細思考資料該如何注入、哪些資訊價值更高。因此聚焦兩件事:第一是資料的收集與注入;第二是 evaluation layer。每次修改都要透過評估確認,變更真的把系統帶向預期方向。兩者必須共存,少一個都不夠,尤其是聲稱能解決 production issues、被放進組織最敏感位置的 AI agent。
So we have to balance between the two and work on 2 on a parallel way.
所以必須平衡資料與評估,並行改善兩者。
Just throwing context at the problem became problematic. What? What were the problems?
單純把 context 全丟進去產生了問題。問題具體是什麼?
There's a lot of context to it to understand, like what are we aiming to do to resolve issues in a production scale? Large organizations, if an SMB's, their scale is very large, could contain a lot of services and one incident can go through maybe dozens of different services to understand how it goes. And when an SRE actually by themselves are investigating the issues, they are looking at so many logs, so many metrics, traces, logs, timelines of stuff, the gate, what has changed, what the new version includes, so many different stuff. Even before, they are so too much context to handle. And a lot of the time there's so much context in a realm that they are unaware of maybe the code themselves, the business plan that connected to it. So they had to involve other people in the organization as well, the product manager, the software engineer, the devils that was responsible at the same day. So getting all this context was very problematic. So we saw one, collecting all this information is a problem. Second is to understand what's critical and what's not, so how to do it and what's not. And we found that the best method is giving the the liberty to do a lot of stuff.
在大型組織,甚至規模較大的 SMB,一次 incident 可能橫跨數十個服務。SRE 調查時要查看大量 logs、metrics、traces、timelines、Git 變更和新版內容,context 本來就多到難以處理;其中還包括他不熟悉的程式碼或相關商業計畫,必須請 product manager、software engineer、當班 DevOps 等人協助。第一個難題是收集全部資訊,第二是判斷哪些關鍵、哪些不重要。我們發現最好的方法是讓 agent 在許多事情上有自由。
But in certain areas, we don't give it the liberty. We say these logs are important. I'm not asking you for two reasons. One is I don't want to have like the logs that happened during this time of the incident. I feel that they are crucial to understanding what the issue really was. That's one. So I don't want that maybe in 1% of the cases it won't be using it. That's first and second, also to save time and tokens, because if I'm giving the AI the liberty to do like the calls that it wishes, so it will do the first round of LLM iteration. Understand that it needs to go through the logs, then call the logs, get it back while injecting it in the first place before the LLM starts. Saves a lot of both time and tokens.
但有些領域不能完全自由,例如我們會明確指定 incident 發生期間的 logs 必須使用,因為它們是理解問題的關鍵,我不想承擔那 1% 被忽略的情況。預先注入也能省時間與 token;若完全讓 AI 自己決定,它得先跑一輪 LLM 才發現需要 logs,再呼叫工具取回。若一開始就放進去,兩者都省。
That's a big thing I've heard from people. It's like AI can process a lot of data, taking a lot of context, but you're paying a ton for every session just in context. And I wonder how you think about limiting that context, because when I talk to people about just logs, not even like all the other data you talked about, logs grow to be massive, massive amounts of data. So how do you sort of narrow that down?
AI 能處理很多資料與 context,但每個 session 光背景資料就很昂貴。Logs 本身就會膨脹成龐大資料,更不用說其他來源;你們怎麼縮小範圍?
So we have a few methods of doing it, but I would say the main methods is that we do is that we like to delegate. So what we're doing, instead of like having our main agent be the one collecting all the information, digesting it, understanding what's going on, and continue with that context as we go through the session, it's just too difficult, it's too big and it doesn't make like a lot of sense. So we have all these very like clips and sub agents that are very good at doing a much smaller task, specifically in the realms of like a log summarization and stuff like that. So they get a lot of context from the main agent, like what is the issue and what kind of investigation do they need to do? It's very important because if I would take like a generic log summarization agent, it doesn't have the right context in order what should I summarize, what should I look at, etcetera, etcetera. So now we have the sub agent that is very good of getting the context from the main agent, very good at doing the summarization and bring back all the summary of the result. And this makes our main agent much more efficient, much faster, less tokens, and less chances to hallucinations.
主要做法是 delegation。若 main agent 自己收集、消化所有資訊,還把這些 context 帶完整個 session,資料既太大也不合理。因此我們使用專門執行小任務的 sub-agents,例如 log summarization。Main agent 先告訴它問題和調查目標;這很重要,因為通用摘要 agent 不知道該看什麼。專用 sub-agent 接收適量任務背景、完成摘要,再把結果帶回,讓 main agent 更有效率、更快、更省 token,也較不容易 hallucinate。
I've talked to folks who, you know, they're sort of ongoing argument about whether it's you have one central generalist agent or you have the the multiple sub agents. And I wonder how you thought about that problem, deciding on having the separate agents.
業界一直爭論應該有一個中央 generalist agent,還是多個 sub-agents。你們為什麼選擇拆分?
So actually it took us quite some time to get to the separate agents because first of all, the foundation models are are really good, right? We at Komodor are not building our own models. We're using off the shelf models, especially Claude and they're doing a very good job. The context window is getting bigger and bigger. But even if the models are capable of doing a lot of stuff, we really want it to be efficient and fast and we don't want to rely on it. We don't want to rely that the model is good enough with handling of hundreds of thousands of tokens, even if it does have this capability, because it will succeed when providing them a few, let's say 10 or 20 thousands of tokens, it will be right 99.9% of the time and when we provided the 200,000, it will be right 98% of the time. 98 is just not good enough. You can say I'm wrong in 49 out of 50 production incidents and it will be right in 50 out of 50. So we must aspire to that 100%. That's why it's very important for us, even when it's under the context window, we still have like a lot of like improvements and optimizations that we are doing.
我們花了一段時間才走到多 agent。現成 foundation models 已經很好,Komodor 不訓練自己的模型,而是使用 Claude 等模型,context window 也越來越大。但能力足夠不代表效率與可靠性足夠。我們不想依賴模型能正確處理數十萬 token:一兩萬 token 可能有 99.9% 準確度,二十萬 token 也許掉到 98%;在 production incidents 裡,98% 遠遠不夠。你不能說五十次有四十九次答錯、一次答對也算可以,我們必須追求接近 100%,所以即使沒超出 context window,仍要持續最佳化。
Yeah, for those sub agents, do you use different models, use cheaper models to do some of the big data stuff?
這些 sub-agents 會使用不同或較便宜的模型,來處理大量資料嗎?
Yeah, yeah. So it depends on the task. If it's a task that we feel that doesn't require a lot of heavy lifting or something like that, we try it out with different models. We do it a lot like trying out different models I talked earlier about like our evaluation layer. So every change that we are making getting into some kind of testing period in which it's being test in production, real life incidents that our users are currently working with our production agents and sub agents. But under the hood we are also running the staging or the non production agents. And this way we can compare the results. We have like a judges that are very good at comparing which one was better and after a few thousands or 10s of thousands depending how much confident we want to gain, which is like OK, the actually the testing sub agent was actually behaving better or behave this with the same accuracy, but we spend much less tokens on it. So we have a lot of things to compare with. Accuracy is our only great of course. And then we select which one was better to us. And like, this is how we keep improving.
視任務而定。若不需要重度推理,我們會試不同模型,並利用 evaluation layer 比較。每個變更都要進測試期:使用者照常跑 production agent,底層同時讓 staging/non-production agent 處理真實 incident,再由 judges 比較結果。累積數千到數萬次後,才能判斷測試版是否更準,或準確度相同但 token 更少。Accuracy 當然是主要標準,再選出更合適的版本,持續改善。
For judging what's right, how do you guide those LLM judges?
要判斷何者正確,你們如何引導 LLM judges?
I don't have like a magic hint here. Something that would work, I think with working with LLM, this is not the only the most important part of your work. I think it's having good judges. Is the base for everything in the pyramid of velocity like the judge or the baseline? If you have a judge that you trust on that, you say yes. I would pick the same thing. I would judge it the same way in 99% of the cases this way, like it improves velocity a lot. So we started off actually with manual evaluations. We built our own judges and then in a separate manner we started like going over them manually, which did a really good job. We felt like the judges are very close but not enough. So we optimize them, optimize them, optimize them and then we just re evaluated. And the easier trick that judges can have is that the input is static. The input, we say this was the issue, this was the way of investigation. Was it good? Did it find the correct root cause? Yes or no? So their input is static, unlike production incidents where the input is dynamic.
沒有魔法提示。建立可信任的 judges 是 LLM 工作最重要的基礎,也是速度金字塔的底層。若 judge 在 99% 情況都會做出和你相同的判斷,開發速度就會大幅提高。我們先人工評估,自行建立 judges,再逐一人工核對,反覆最佳化與重新評估。Judge 的優勢是 input 固定:問題是什麼、調查如何進行、是否找對 root cause,都能保存後重複比較;production incident 的 input 則會隨 logs、metrics 和 services 持續改變。
It's not that I can investigate like 2 hours later because everything has changed, the metrics changed, the logs changed in new services came up, board were crashed so everything changed. So in this case of the judges the once the input is static I can improve it as we go. And also cool trick that we have find that judges are way better than saying and much more consistent. When I'm asking them who was better A or B, they are much better in with me saying just give me a score for B or for A, they will say sometimes 70758085. So when I'm saying who was better A&B, they are much more consistent. So when we are doing this AB testing, we get a really good confidence and really which was better and we are very confident about like selecting whoever and also that the number of investigation is also important here. So after 10 investigations I wouldn't feel confident enough even in all 10B1A. Even after 10 I wouldn't say like I have confidence to say that like B is better. I need it to be at least a few thousands, maybe 10s of thousands. Depends on the change.
Incident 過兩小時再調查,環境可能全變了;judge 的固定 input 卻可持續改善。我們也發現,相較要求它替 A 或 B 各自打 70、75、80 分,直接問『A 和 B 誰比較好?』結果更一致。這讓 A/B testing 更可信,但樣本量仍重要:即使十次都是 B 勝,我也不會斷言 B 更好;至少需要數千、甚至數萬次,取決於變更大小。
Do you build new training incidents off of real incidents? Like if an agent sort of doesn't quite get it, do you like retrain it, retrain it on it.
你們會根據真實 incident 建立新的訓練案例嗎?如果 agent 沒處理好,會用那個案例重新訓練嗎?
Yeah. So we can't really use our customers for data for this, but we do have a way of like anonymizing data of taking the the essence of an issue that we saw and turn it over something we call a scenario. Scenario is something like we have like own side project, side repos and it is something we know how to deploy. First I'm deploying 2 namespaces with two different parts, a database, a Redis, etc. Connected to some kind of an AWS service and so all this cloud infrastructure that is talking together. We include a lot of noise just to simulate production environment. We include a lot of services that are unrelated to the issue that are very noisy. We include a lot of errors and alerts that are not related to the issue that we want to test just to make sure we simulated the world in a proper way. And then we break it. We do the change that caused the break or starts off with a break, it depends. And then we do our own agents and we validate it and it also gives us some kind of regression testing. We are on a daily on all of our scenarios and making sure we are not now. I don't know, something happened and we forgot about our postal sub agent and is doing less of a job.
我們不能直接拿客戶資料訓練,但能匿名化並擷取問題本質,轉成稱為 scenario 的測試案例。我們有自己的 side projects 和 repositories,可部署兩個 namespaces、資料庫、Redis、AWS service 等互連的 cloud infrastructure;也會加入大量無關服務、errors、alerts 和 noise,模擬真實 production。接著刻意破壞系統或套用會造成故障的變更,讓 agents 調查並驗證。所有 scenarios 每天執行,形成 regression testing,確保某個 sub-agent 不會因其他變更而退步。
So we want to make sure that we are good not only for today but also for tomorrow.
我們不只要今天表現好,也要確保明天仍然可靠。
The other half of SRE is sort of finding and validating a fix and and the other half of your talk was about self healing SRE agents. You know, with all this judges and how do you get to a confidence space where you can let the agent just push fixes to incidents?
SRE 的另一半是找出並驗證修復方法,你的演講也談到 self-healing SRE agents。要累積到什麼程度的信心,才能讓 agent 直接把修復推到 incident?
I can tell you what we found out up until now. First, we do see a lot of appetite to it like we do see that a lot of our customers are actively asking about the we do have this capability of doing it like autonomous way. But having said of that, we don't see a lot of like like production environments that are fully automated with like AI doing the investigation, offering the remediation and actually performing the remediation offered all in this together. What we do provide here at Komodor is we do all of that, but with human in the loop, but with someone that Claudia, our AI agent, that's her name, she's doing the investigation later on she's doing the remediation plan. The remediation plan ends with this is the command needs to be executed. These are the reasons why I think this is the command needs to be executed with all related evidences, etcetera. And then the users can get in the way and say like whether they approve or reject the suggested fix. And if they approve, Claudia does that ourselves. We do see some time that users are not proving or rejecting, but copy pasting it and doing it themselves. We also consider that as the win.
客戶對 autonomous remediation 很有興趣,但真正完全自動化的 production environment 還不多:讓 AI 完成調查、提出 remediation 並直接執行,大家仍很謹慎。Komodor 的 Claudia agent 可以完成整個流程,但保留 human in the loop。她先調查,再提出 remediation plan,列出要執行的 command、理由和證據,使用者可以核准或拒絕;核准後才由 Claudia 執行。有些人不按核准,而是複製 command 自己操作,我們也視為成功,因為他相信修復方向正確。
Yeah, if they thought that our remediation was the correct 1. So we consider it as a win. So we do see the world where more appetite to it. We do feel there's like foot in the door that is there. We do see like a lot of testing playgrounds, etc that are using like the autonomous remediation just to gain more trust in it. We do see like in certain areas where we are being asked and we do have the capability of allowing autonomous remediation, but only for let's say a rollback only. Let's say for certain actions that I'm confident that I do not like create a new version or something like that or apply the code fix, let me do it myself. And stuff that are more like soft. They have like less issues to adopt. And we do see like the trend and the foot in the door getting there. But as I, as I started off, we do ask ourselves the same question of how can we gain the trust more quickly? What do we need to do? And we feel like the road is clear.
市場對自動化的接受度正在增加。許多人先在 testing playgrounds 使用 autonomous remediation 建立信任,也有人只允許特定低風險動作自動執行,例如 rollback;涉及新版本或 code fix 則保留人工處理。這些較柔性的操作更容易採用,算是先踏進門內。我們一直問如何更快取得信任,而前進路線已逐漸清楚。
Yeah, I mean, the the trust issue is a just a huge thing for AI right now. But I think anything that minimizes the time SRE spends at 2:00 AM during an incident, it's good, right?
Trust 是目前 AI 的巨大問題。不過只要能減少 SRE 凌晨兩點處理 incident 的時間,就是好事。
Yeah, definitely. Both from the SRE perspective, both of the system perspective, right? Time is money, and if I can be down not for 15 minutes but for 30 seconds, that means a lot.
當然,無論對 SRE 或系統都是如此。Time is money;服務中斷 30 秒,而不是 15 分鐘,差別非常大。
My man's still an accountant, huh?
你心裡果然還是個會計師。
I wonder, you know you have a lot of incident experience with customers. I wonder if if doing the AISRA tool has pinpointed a part of infrastructure that is particularly fragile.
你看過許多客戶 incident。開發 AI SRE 工具後,有沒有發現哪一類 infrastructure 特別脆弱?
I do think that the next line of stuff that we started off with detection, we moved off to remediation and saw it to investigation. Then we, we are now working like on remediation and gaining trust and what we talked about. And I don't think like the next cycle of things would be much more preventive. And in the world of preventive, you need to come up with like what's the main issues that are currently taking place, what's the most common stuff, what's the problematic areas in my infrastructure, etcetera. So we do have some agents that are capable of doing a lot of things. And Claudia has a sub agent that is in the preventive side of things that you can talk to Claudia, she will then activate a preventive agent. It will go through your infrastructure and imply also we have the possibility of investigating and already resolved the issue. And why is that? Because we see a lot of people coming back to resolve the issue to look what it was and they talk with AI, how did it happen? How can I improve in the future? So Claudia is very good in in helping preventing stuff from the long term.
我們的發展先從 detection 到 investigation,現在處理 remediation 與信任,下一循環會更偏 preventive。預防工作要找出常見問題和 infrastructure 的薄弱區。Claudia 已有 preventive sub-agent,可檢查 infrastructure,也能回頭調查已解決的 incident。許多人修完後仍會問 AI:事情怎麼發生、未來如何改善。Claudia 很擅長協助長期預防。
So she can see maybe if someone resolved an issue by applying a very specific and tailored patch to a certain area, so she would offer that. So it's good that you catch this specific thing, but the real root cause isn't the code. And this is the place where we need to go to investigate. Would you like me to do it for you? So this brings a lot of maturity to the product and brings like a lot of trust because we do see a lot of those like we got a bit surprised that we see how many people are going back to issues that were. Actually resolves going back and start to to come up with a proper way on how to prevent stuff like this in the future. And going back to your original question of do we see. So yeah, we do see a lot of stuff that are pretty common or repetitive especially in the world where you manage like Helm or Argo-style deployments and this is the way how you do like git OPS, et cetera. So we do see in this earned DB related issues, connection related issue, networking related issue that are happening a lot. And once we see an issue that is becoming very repetitive, we also consider creating a sub agent. We do the A/B testing that we started talking about it.
例如有人只在局部套一個非常客製化的 patch,Claudia 會指出:你處理了表面問題,但真正 root cause 可能不在這段 code,並詢問是否要繼續調查。這能提高產品成熟度與信任。我們意外發現很多人會回頭研究已解決的問題,尋找未來預防方法。常見重複問題包括 Helm/Argo 與 GitOps deployments、database connections 和 networking。當某類問題反覆出現,就考慮建立專門 sub-agent,並用前述 A/B testing 驗證。
So once we feel that there is more like specific pinpointed speciality that needs to be applied, invoke our sub agents to it and they're like they're doing a pretty decent job.
只要確認某個領域需要高度專門能力,就叫用相應 sub-agent;目前效果相當不錯。
It makes sense to me that people are going back to solve issues to see what would happen. Like that's the the post mortem process, right? You want to build some playbooks and learn from it. Do you have any sort of like automated postmortem playbook generation based on previous instance?
回頭研究問題就是 postmortem:建立 playbooks 並從中學習。你們能根據過往 incidents 自動產生 postmortem 或 playbook 嗎?
Yeah. So this is something we do have in like 2 separate ways. 1 users are capable of like uploading and connecting us to their like knowledge databases or some kind we can connect to it, read their past postmortems and improve our investigations going forward. Second, we have like a memory layer which like investigations that are taking place that we can learn from future investigations. So we definitely do this. And also when you like trigger Claudia, once an issue has been resolved, you can talk to her, ask her anything she would like. She has all the context. Of course, she's familiar with the issue that just got resolved or got resolved yesterday and she understand why it happened and how did it got there, etcetera, etcetera. It brings us to the point where she's very good at like not only doing the investigation, but also like closing the loop and providing a suggested postmortem that you can like work on and refine on it and make it better. We see the world that people are like having products such as Claudia off the shelf, like AI SRE, but we see a lot of people starting like building their own orchestrator agent, like using Claudia not as the interface but as the brain.
有兩種方式。第一,使用者可連接自己的 knowledge databases,讓我們閱讀過去 postmortems,改善之後的調查。第二,我們有 memory layer,從歷次 investigations 中學習。Incident 解決後,使用者仍可詢問 Claudia;她保留完整 context,了解問題如何發生,能協助閉合迴圈、產生 postmortem 草稿,再由人調整。我們也看到企業建立自己的 orchestrator agent,把 Claudia 當作 brain 而非 interface,讓它和其他 agents 協作。
So they connect, they build their own like orchestrating agent which is capable of calling other agents such as ours. And I think this is a very interesting world and I think this is where the world is gonna be. We say like 6 months ago, 100% of the people would just go to our UI and this is what 100% of the usage of Claudia was through our UI. And like in the past three months, four months, something like that, I can say it's changes because it's still like 80% of the usage. But we do see like this higher usage of being from our MCP, from our API, from Slack, like people not even going into using their own interfaces, agents, whatever. In some cases, they don't even know that Claudia was being triggered behind the scenes because it happened in like somewhere agentic flow somewhere.
六個月前,Claudia 幾乎 100% 從我們的 UI 使用;近三四個月已降到約 80%。越來越多呼叫來自 MCP、API、Slack 或客戶自己的介面與 agents。有時使用者甚至不知道 Claudia 在後台被觸發,因為她只是更大 agentic flow 的一環。這很可能就是未來。
I mean that that makes sense, right?
這很合理。
Developers, software engineers love to build stuff. There's a sort of conversation going on right now is like about the shift in build versus buy, where people are like, why don't I build everything? And I think people are kind of realizing again why they don't build everything right.
Developers 很愛自己造工具。現在又出現 build versus buy 的討論:『為什麼不全部自己做?』但大家也開始重新想起,當初為什麼不會什麼都自建。
I think like the era of open source I think is coming back in. Some people say I can use an open source for it or I'll build something myself. I don't need this like off the shelf something. And we do see like it's getting really good result and a really good demo. It creates a really good demo when you're tackling production incident with a lot of noise and like real life happens. You want to get the professional tool and not the tool that I just built yesterday in my garage.
Open source 的時代似乎又回來了。有人會說,拿開源方案或自己做就好,不需要 off-the-shelf 產品。這確實很容易做出漂亮 demo;但面對充滿 noise 的真實 production incident,還是會想要專業工具,而不是昨天才在車庫做出來的東西。
Yeah, there's no support for that one.
而且車庫版本沒有技術支援。
So you have a pretty good view on what SRE is. What do you think the SRE of the future will be doing?
你很了解 SRE。你認為未來的 SRE 會做什麼?
So I think they will be keep doing resolving issues, but I think they will put a lot of more focus on like cost optimization and planning ahead and doing the strategic stuff that they never get the chance to do because they are putting down fires 90% of the time. They are always in stress. And I don't think they're going to move to Haiti and now drink the pineapple juice. I think they will keep working hard and they will have a lot of work to do. But I think there's a lot of stuff that investigation is just too difficult. There's a lot of stuff where the investigation requires too much expertise and too many people involved. So in this world like I think AI will lead the investigation and the ISRE will be like some kind of the manager as we see with the world of agents where all people become managers of their own agents. So I think AISRAS will manage other people with agents and their own agents and the reviewers, the approvers of workflows and let like the agents do the role of like putting much less tension on their shoulders of the SRE like it is exists today.
他們仍會解決問題,但會投入更多時間在 cost optimization、前瞻規劃和策略工作;現在 90% 時間都在救火,根本沒有機會做這些事。他們不會搬去海地喝鳳梨汁,仍然會很忙。只是有些 investigation 太困難,需要太多專家參與。未來 AI 會主導調查,SRE 則像 agent manager:管理自己的 agents 與其他人,擔任 reviewer 和 workflow approver,讓 agents 減輕目前壓在 SRE 肩上的負擔。
It's that time of the show where we shout out somebody who came under Stack Overflow, dropped some knowledge, shared some curiosity and earned themselves a badge.
現在來表揚一位在 Stack Overflow 分享知識與好奇心、並獲得 badge 的社群成員。
Today we're shouting out the winner of a populist badge, somebody who dropped an answer that was so good it outscored the accepted answer. Congrats to Joe Mahler for answering. Is there a Python equivalent to the which command? If you're curious about that, we'll have the answer for you in the show notes. I'm Ryan Donovan. I edit the blog, host the podcast here at Stack Overflow. If you have questions, comments, concerns, topics to cover, e-mail me at podcasts@stackoverflow.com. And if you want to reach out to me directly, you can find me on LinkedIn.
今天表揚 Populist badge 得主 Joe Maller;他的回答甚至超越被接受答案。問題是:『Python 有沒有相當於 which command 的做法?』答案會放在 show notes。我是 Ryan Donovan,負責編輯 Stack Overflow blog 並主持 podcast;問題或建議可寄到 podcasts@stackoverflow.com,也可在 LinkedIn 找到我。
My name is Asaf Savich. I'm AI engineering group manager at Komodor. You can learn more about Komodor and it's AI agent Claudia by going to Komodor dot IO. And if you want to link me, feel free to add me in LinkedIn by searching me Asaf Savich ASAFSAVICH.
我是 Asaf Savich,Komodor 的 AI Engineering Group Manager。想了解 Komodor 與 AI agent Claudia,可前往 Komodor.io;也歡迎在 LinkedIn 搜尋 Asaf Savich 與我聯絡。
Thank you for listening everyone, and we'll talk to you next time.
謝謝大家收聽,我們下次見。