← 回到所有學習包
Software Engineering Daily · 48:28

SED News: The Kimi Moment, Runaway AI, and Tokenmaxxing

本集由 Gregor Vand 與 Sean Falconer 討論失控 AI 事件、代理成本失控、Kimi 與開放權重模型,以及近期開發者社群熱門話題。

開始聆聽

目前章節:尚未播放

時間碼為章節起始位置;逐句稿不偽造 Spotify 未提供的精確時間碼。

章節導讀

CHAPTER 1

開場與每月科技新聞巡覽

Episode Introduction and Monthly Tech Headlines Overview

兩位主持人從旅行、家庭與機上 Starlink 聊起,並介紹本集將涵蓋失控 AI、Kimi 與 Hacker News 熱門話題。

CHAPTER 2

Anthropic、OpenAI 與外部組織存取事件

Anthropic and OpenAI Models Accessing External Organizations

Claude 與 OpenAI 模型在網路攻防評估中碰觸外部系統;主持人認為事件核心多半是隔離、監控與 human-in-the-loop 的人為失誤。

CHAPTER 3

糟糕 Agent Loops 與災難性超支

Catastrophic Budget Overruns Blamed on Bad AI Agent Loops

Amazon 的 AI 預算五個月超支約 860%;沒有明顯崩潰的 agent loop 會安靜地燒掉 token,凸顯 observability、預算上限和成本預估的重要性。

CHAPTER 4

Microsoft 的 AI 漲勢與 Apple 五兆美元時刻

Microsoft's AI-Driven Surge and Apple's Trillion-Dollar Mark

Microsoft 以可替換模型的基礎設施降低單一供應商依賴;Apple 則靠 iPhone、服務與品牌忠誠維持高估值,同時面臨記憶體供應問題。

CHAPTER 5

Waymo 與 Uber 的自駕車合作走向拆夥

Autonomous Vehicle Partnership Dissolves in Key US Cities

Uber 曾替 Waymo 補足需求與車隊營運,但 Waymo 成長後希望端到端掌控服務,雙方在美國數個城市的合作因而逐步瓦解。

CHAPTER 6

Ray 創建者 Anyscale 加入 Neo-cloud Nscale

Ray Creator AnyScale Joins Neo Cloud N Scale

Anyscale 的 Ray 與受管服務被 GPU-as-a-service 公司 Nscale 收購;討論也延伸到 Scale AI、Meta 與 Fireworks 的人事和資本動向。

CHAPTER 7

Moonshot AI 的 Kimi 挑戰 Frontier Models

Moonshot AI's Kimi Model Challenges Frontier AI

DeepSeek 先改變市場預期,Kimi 再以高速迭代、開放權重與亮眼的 MCP benchmark 追近閉源模型,同時引發評測偏差與企業治理疑問。

CHAPTER 8

晶片禁運如何逼出中國實驗室的創新

How Chip Embargoes Push Chinese Labs to Innovate

資源稀缺迫使中國 AI 團隊提高效率,就像早期 Google 以有限硬體創新;Kimi K3 公開完整權重,也讓 model provenance 更具商業價值。

CHAPTER 9

模型多元化與 Meta 的 Llama 路線

Diversification of Models and Meta's Llama Approach

企業傾向同時採用開放權重與閉源模型,以降低禁令、價格和供應商風險;Meta 雖早押注 Llama,卻可能在執行上錯失先機。

CHAPTER 10

JetBrains、重構、CodePen 與 DoomQL

JetBrains, Refactoring, CodePen, Doom QL, and Future AI Trends

Hacker News 精選包括穴居人語氣的 token 實測、用重構降低 input token、CodePen 2.0 即時協作,以及用 SQL query 渲染 Doom。

核心中英對照

時間碼會跳到該句所屬章節;中文以傳意為主,技術詞保留英文。

There is a part of this that is not really an AI problem; it's the boring problem of a human not being in the loop at the right time.

這有一部分根本不是 AI 問題,而是很無聊的老問題:人類沒有在正確時間進入流程。

It just feels like we haven't thought through all those controls for AI right now.

現在看來,我們還沒有替 AI 想清楚整套控制機制。

Every model is substitutable.

每一個模型都可以被替換。

We've gotten used to the idea that these Chinese lab open weight models are competitive.

我們已經習慣中國實驗室的開放權重模型具有競爭力。

There's a certain scarcity of resources that they have to deal with.

他們必須面對某種程度的資源稀缺。

Good design leads to optimized token cost.

良好設計也會帶來更低的 token 成本。

本集必備詞彙

runaway AI
表面看似脫離控制、自行採取危險行動的 AI;本集強調許多案例其實源自人為設定與監控失誤
guardrail
限制模型權限、行為或輸出的安全控制;也可能同時妨礙合法調查工作
human in the loop
在關鍵節點由人類監督、審核或中止自動化流程
agentic loop
代理反覆規劃、呼叫模型或工具並修正結果的迴圈;若失控可能持續消耗 token 而不當機
token-maxxing
把大量使用 token 當成目標或績效,卻未確認是否產生相應價值
observability
透過記錄、指標、追蹤與警報理解系統內部正在發生什麼
open-weight model
公開模型權重、可由使用者自行部署或檢查的模型;不必然等同完整開源
model provenance
模型及其輸出來源、訓練與版本脈絡的可追溯性
model distillation
用較強模型的知識或輸出訓練較小模型,使後者更快取得接近的能力
neo-cloud
主要提供 GPU 等 AI 運算資源,而非全套一般雲端服務的基礎設施供應商
Goodhart's law
當某項指標成為目標,人們會針對指標最佳化,使它不再能可靠衡量原本目的
input token cost
模型讀取提示、程式碼與背景資料所產生的 token 費用;良好重構可減少必要 context

理解題

1. 為什麼主持人認為多數『失控 AI』事件仍是人類控制問題?

人類先賦予任務與網路權限,卻未做好隔離、重複行為監控、failsafe 或及時的 human-in-the-loop;模型並非自行逃出限制。

2. Agent loop 的成本風險和傳統無限迴圈有何不同?

傳統無限迴圈通常會當機或觸發明顯警報;agent loop 可持續產生看似合理的結果、不斷呼叫模型並悄悄累積帳單。

3. Microsoft 為什麼強調『每個模型都可替換』?

企業同時使用多種模型可降低對單一供應商的依賴,保留切換技術方案與議價的能力。

4. DeepSeek 與 Kimi 對市場造成的心理衝擊有何不同?

DeepSeek 首次證明中國開放權重模型能接近 frontier labs,造成巨大震撼;Kimi 同樣驚人,但市場已習慣這類競爭力,因此反應較平靜。

5. 晶片稀缺為何可能反而促進創新?

團隊無法單靠昂貴硬體解決問題,就必須提高架構、訓練與部署效率;早期 Google 也曾因有限資源發展出重要的基礎設施創新。

6. 重構如何降低 AI coding agent 的 input token 成本?

良好模組化與 DRY 設計讓 agent 能辨識並只讀取相關的小型檔案集合,而不必把大量無關程式碼塞進 context。

完整雙語逐字稿

英文以 Spotify 自動逐字稿為主,並以 Software Engineering Daily 官方頁面確認主持人、音檔與可確定的技術名稱。Spotify 不保證自動稿準確;標記片段仍可能需要對照原音。

1. 開場與每月科技新聞巡覽Episode Introduction and Monthly Tech Headlines Overview
Gregor Vand

Hello and welcome to SED News. As I think many of you know by now, this is the monthly format of Software Engineering Daily where we dive into the tech headlines. We go into a deeper topic in the middle and then we just do a spin around our favourites from Hacker News, Highlights and Spoiler. There's also a fun one that didn't appear in Hacker News, but we'll get to that at the end as well. But yeah, as usual, I think I saw you last, Sean, as opposed to spoke to you, saw you in Singapore, which was fun. Yeah. So you've been travelling as you often are, but travelling in my neck of the woods was fun to see you.

歡迎收聽 SED News。這是 Software Engineering Daily 的每月新聞節目,我們會深入討論一個主題,再快速巡覽 Hacker News 上各自喜歡的焦點與彩蛋;最後還有一則沒上 Hacker News 的趣聞。Sean,我們上次不是通話,而是在新加坡碰面,真的很有意思。

Sean Falconer

Yeah, it was great. It was my first trip there, so it was great to hang out. And that's two times in like a few months. This might become a regular thing. We might just have to start doing these in person.

對,那是我第一次去新加坡,很高興能一起聚聚。短短幾個月就見了兩次,說不定會變成固定行程,以後乾脆面對面錄音。

Gregor Vand

Yeah, and Spoiler and coming back to SF in a couple months. So yeah, it's good.

對,而且我幾個月後也會回舊金山,所以很好。

Sean Falconer

And I'm off to not quite Singapore, but I am off to Australia here in the next week.

我下週也要出發,雖然不是新加坡,但會去澳洲。

Gregor Vand

Nice, but yeah, what else has been keeping you busy over the month?

不錯。這一個月還有什麼事讓你忙個不停?

Sean Falconer

I mean, a summer I feel like has just flown by. Like I feel like my kids ride a school and then suddenly it's like August and they're going to be back to school in a few weeks. So things have become really hot and fast this summer. I think we've done a lot of traveling and then just stuff moves quickly. But been a big summer for my, I usually don't talk that much about my family, but big summer for my son who learned to swim, learned to ride a bike, and also now has gotten significantly better at reading. So I'm very proud of the amount of work that he's put into this summer to learning new skills. How about you?

我覺得暑假一下就過去了。孩子才剛放假,轉眼已經八月,再過幾週就要開學;這個夏天又熱又快,我們旅行不少,所有事情都動得很快。我平常不太談家人,但兒子今年暑假學會游泳、騎腳踏車,閱讀也進步很多。我很為他投入學習新技能的努力感到驕傲。你呢?

Gregor Vand

That's like a whole model change there for your son.

你兒子這根本是完成了一次完整的模型升級。

Sean Falconer

Yeah, it's the new Kimi model, Yeah.

對,這是全新的 Kimi 模型。

Gregor Vand

Yeah, on my side just, yeah, a bit of travelling and I'm recording this on my side from Scotland. I try and come back here a couple times a year. So yeah, nice to get out of the city. I'm up in the the Highlands of Scotland, so lots of nature around. Yeah, as I was flying just back to something we've talked about a few times, Starlink on planes, but yeah, that used to be super seamless with like no login screens. And that was like, I believe dictated by Starlink, but they've apparently had to cave into airlines wanting to put, like, a login interstitial for that. So yeah, I mean, it's such a trivial thing, but like, expecting starting to just connect. And then suddenly I found the airline saying, hey, if you need to log in with your membership number. And I thought, what's happened here? And I read up on it, Yeah.

我這邊也旅行了一些,現在人在蘇格蘭錄音。我一年會回來幾次,目前在高地,能離開城市、親近自然很舒服。說到飛行,我們談過幾次飛機上的 Starlink:以前完全不用登入,非常順暢,據說是 Starlink 的要求;但他們似乎向航空公司妥協,開始加入登入頁。雖然只是小事,但原本期待一連上就能用,突然被要求輸入會員號碼,讓我很疑惑。

Sean Falconer

Do you still like, do you have to be up in the air or does it work immediately?

它一定要起飛後才有訊號,還是在地面就能用?

Gregor Vand

That's a good question. For some reason, I don't think it was working kind of gate to gate on the flight I was on. This was Qatar Airlines again. Yeah, but I think it's still supposed to. But yeah, still, you just have to log in with your Qatar membership number now, which you don't need to be a special tier or anything. You just literally have a number so.

好問題。我搭的那班 Qatar Airways 似乎沒有做到從登機門到登機門都能使用,雖然理論上應該可以。現在只要用 Qatar 的會員號碼登入,不必有特殊會籍,單純有個號碼就行。

Sean Falconer

It seems like United is doing some experimentation, but I think I've only had like one or two times I got like a text message saying that there's Starlink on this flight. And then I think both times they ended up having problems with the plane and having to switch planes and then we lost Starlink.

United 似乎也在試驗。我只有一兩次收到簡訊說班機有 Starlink,結果兩次飛機都出了狀況、換機後就沒有 Starlink 了。

Gregor Vand

Yeah, that's the one thing about Qatar. They really rolled out like big time across most of their fleet now. So if you're on a Airbus A350, you know you'll have it, which is always nice. But right, out of plane chat, which I can easily head into at any time and onto software. So onto the headlines. Yeah.

Qatar 的優點是已大規模部署到多數機隊;搭 Airbus A350 幾乎可以確定有網路。好了,我很容易一直聊飛機,現在回到軟體與新聞標題。

2. Anthropic、OpenAI 與外部組織存取事件Anthropic and OpenAI Models Accessing External Organizations
Gregor Vand

So there's been a sort of flurry of, we're just sort of calling it runaway AI stories over the last couple days. And the first one is actually the most recent. So that's just disclosed today or yesterday, which was Claude. And this is actually in reaction to OpenAI, which we'll also talk about, but basically in reaction to something that OpenAI disclosed, Anthropic has now also disclosed that it's that Claude hacked into 3 organizations whilst they were just testing their cyber capabilities, or at least that's what they said. And they basically said that Claude had gained unauthorized access to outside companies during an evaluation of its cyber offensive tasks. And it was a misunderstanding, apparently, that Claude had access to the Internet in its testing environment. We're going to see this as a kind of theme here of, like, is this really a runaway model, or is it just a human who, like, forgot to do something? But yeah, I mean, this one feels like human just forgot to place it under strict, no Internet access controls. But yeah, what did you think of this one?

最近幾天冒出一連串我們稱為『失控 AI』的新聞。最新一則是 Claude;在 OpenAI 披露類似事件後,Anthropic 也承認 Claude 在測試網路攻擊能力時,未經授權進入三個外部組織。官方說法是測試環境誤讓 Claude 連上網際網路。這會成為今天的主線:真的是模型失控,還是人忘了做好隔離?這一件看來很像人類沒有嚴格關閉網路存取。你怎麼看?

Sean Falconer

Yeah. I mean, I think that all this really ends up coming back to some human decision making, right? Like that perhaps the intention was not to have the model like hack something, but essentially something got lost on the human side of we made a mistake take here and gave it access to the Internet or we made a mistake here or whatever. The issue is, it really comes back to some human level of control and what guardrails in place. So, and then the thing though, with all this that has me thinking from both the OpenAI attack on Hugging Face and then also this latest one with Claude is that if you can just apologize and say, like, oh, it was unintentional, it was an accident and stuff like that. When the stuff happens where there was intention for abuse or misuse or something like that. Like does it give people basically license to excuse the fact that they are doing something potentially malicious and just say, oh, it was an accident? I mean, we've seen that with viruses too.

最後仍要回到人類決策。也許本意不是讓模型入侵,但有人誤給了網路權限,或在其他環節犯錯;核心是人類控制與 guardrails 是否到位。OpenAI 攻擊 Hugging Face 和這次 Claude 事件也讓我思考:如果一句『非故意、只是意外』就能道歉了事,日後真正有惡意的人是否也能拿意外當藉口?這種問題並不新鮮,病毒時代就發生過。

Sean Falconer

I remember when the viruses, I can't remember if it was the one, the back in the early 2000s, there was all these viruses that were like these scripts that went around in emails and then you click on it and we copy your contacts and then fire off to them. And it would take down people's mail servers. And I, I believe like one of those started as a student who was like testing something in like an MIT lab and then it accidentally got out. And I don't know what happened to that student, but like this is certainly not something new. But I just wonder, like, where does it kind of stop with being able to provide an excuse of like this was an accident versus there was intent behind it?

二〇〇〇年代初有許多透過電子郵件散播的腳本病毒;使用者一點開,它就複製聯絡人並寄出,甚至拖垮郵件伺服器。我記得其中一個好像是學生在 MIT 實驗室測試,結果意外流出。這並不是新問題;真正難的是,『意外』和『帶有惡意』之間的責任界線到底在哪裡。

Gregor Vand

Yeah, I'd like just to sort of give context for anyone who hasn't been keeping up with the news. The OpenAI one was the fact that basically, yeah, it hacked Hugging Face. They set it off to do something and it ended up repeatedly trying to get into Hugging Face. But what was the secondary part to that story, courtesy of TechCrunch, was the fact that actually Hugging Face could have detected this like earlier, but they had a sort of human failure on the security end that.

補充背景:OpenAI 的模型是在執行任務時反覆嘗試進入 Hugging Face。TechCrunch 後來指出,Hugging Face 其實可以更早偵測到,只是安全流程也出現人為失誤。

Sean Falconer

Yeah, they caught it, but they just, they didn't escalate it to a human fast enough, so.

他們有抓到訊號,只是沒有及時升級給人類處理。

Gregor Vand

Right, exactly so their systems.

沒錯,也就是說系統其實有運作。

Sean Falconer

Actually work the way it is. So in some ways, like there's a part of this that is not really an AI problem, it's kind of the boring problem of exactly like a human wasn't in the loop at the right time.

所以某種程度上,這根本不是新奇的 AI 問題,而是很無聊的老問題:需要的時候,human in the loop 不在正確位置。

Gregor Vand

Yeah, and I think that's the piece, like it's hit mainstream headlines as like AI is. That's exactly what we all were saying, that AI can be catastrophic, it can go off, it's got a mind of its own, it's going to, it's hacking things left, right, centre. But especially in both cases, the OpenAI and anthropic cases, someone a human did set it up to do something along these lines. But the human did not a keep a watch on just this repeated actions it was taking and it clearly there was no fail safe for it to stop at any time. Again, these are all human determined things that can be set, but it's not like that broke out of the cage exactly. And it seems there's some quite fundamental things that could have been done by a human that just weren't. And that's actually, yeah, it's a quote, boring story, unfortunately, because AI isn't sort of this wild animal quite yet.

主流標題把它寫成 AI 災難性失控、像有自己的意志,到處入侵;但 OpenAI 與 Anthropic 兩案都是人類先設定了相近的任務,之後沒有監看模型反覆採取的行動,也沒有設定能隨時中止的 failsafe。這些都是人能決定的控制。模型並不是自己逃出籠子,而是人沒做幾項很基本的事;所以遺憾的是,這其實是個『無聊』的故事,AI 還不是野獸。

Sean Falconer

Yeah, I mean, you want to sensationalize the headlines, right? Like even my sister who's not in tech at all sent me. The headline was like, this is scary so that you get the attention there. But I think one of the things that was interesting about the hugging face attack was when they tried to investigate, they couldn't actually use Claude or GPT because those models have safety guardrails in place where you can't tell whether are you an incident responder or an attacker. So they have essentially mechanisms in place that like I can't go to Claude until say, like, help me break into the Pentagon or something like that. Like it's going to prevent me from doing that. That means that you also are limited in using those models to investigate. So what they ended up using an open weight model out of China for the forensics. And then I think that also we're going to talk a lot about the open weight models from China later in the episode. But that's interesting consequence of this is that we have these guardrails in place, but there's always ways to I think manipulate the models even with the guardrails in place to do stuff.

新聞當然想聳動,連完全不做科技業的姊姊都傳給我說『這太可怕了』。Hugging Face 案另一個有趣之處是,調查時不能直接用 Claude 或 GPT,因為安全 guardrails 無法判斷你是事件應變人員還是攻擊者;你不能叫 Claude 幫你闖進五角大廈,同樣也可能無法讓它協助鑑識。因此他們改用中國的開放權重模型。稍後我們還會談很多這類模型;這凸顯 guardrails 的兩面性:它可能被繞過,也可能妨礙正當的事件調查。

Sean Falconer

But then when you want to use it for something intentional like incident response, you might not be able to do that because the protections are in place there in the 1st place. And then you have to circumvent it by going to a model where maybe there's less safety guardrails in place.

也就是說,當你真心想把模型用在 incident response 時,原本的保護反而讓你做不了,只能改用安全限制較少的模型。

Gregor Vand

Exactly.

正是如此。

3. 糟糕 Agent Loops 與災難性超支Catastrophic Budget Overruns Blamed on Bad AI Agent Loops
Gregor Vand

And the sort of final one on the runaway story is actually Amazon, not security related, but cost related. We have touched on this the last couple of SED news episodes, just where's the tipping point of cost overruns making it not viable for businesses to be allowing employees to sort of quote Token Max and this kind of thing. But yeah, Amazon basically said that they had a ton of unplanned spend. They called it catastrophically expensive. This is all courtesy of the FT Financial Times. Apparently they had about 860% budget overrun over 5 months and this was basically, in their word, caused by bad agent loops. Just didn't crash loudly enough and they just kept being built. So yeah, I mean, it's pretty interesting for Amazon to come out and actually say that, So yeah.

最後一則『失控』故事來自 Amazon,這次不是資安,而是成本。我們前幾集已談到企業讓員工 token-maxxing,成本何時會高到不可行。據 Financial Times 報導,Amazon 把大量非預期支出稱為『災難性昂貴』:五個月超支約 860%,官方歸因於糟糕的 agent loops。迴圈沒有明顯崩潰,只是不停執行與累積。Amazon 願意公開承認,確實很值得注意。

Sean Falconer

I just don't understand how they could be surprised by this. I feel like we've been beating this drum for months now, and I think this is just the beginning of these kind of stories that we see. But if you build a leader board to encourage people to use AI and that's the metric you're optimizing for, but there's no connection to the value of the use of the AI. Like what do you what do you think's going to happen? Like this is really Goodhart's law essentially showing up on some sort of schedule. Like you're rewarding people for tokens consumption. So it's like you're giving people a license to be wasteful and not looking at like the productivity metrics of those. So that's like the epitome of token maxing.

我不懂他們怎會意外,我們已經敲了好幾個月警鐘,這大概只是開端。如果設排行榜鼓勵大家多用 AI,唯一 KPI 是使用量,卻不衡量使用創造的價值,那結果還能是什麼?這就是 Goodhart's law:獎勵 token 消耗,等於發給員工浪費許可證,卻不看生產力,正是 token-maxxing 的極致。

Gregor Vand

So yeah, they kind of have themselves to blame in that sense, but yeah.

從這個角度看,他們確實只能怪自己。

Sean Falconer

It's kind of similar to like the initial stories we were talking about with like hugging faces. Like, it's still not AI necessarily, just like running rampant. It's a human decision that in both cases, even though like 1 is this attack vector and the other is spend, but it's still like a person making the decision at Amazon to say like, hey, we're going to just have a KPI where we're just going to reward people for maxing out tokens. The other thing that's interesting about this that people have to think about is that with traditional software, when you have like a bad for a loop or code that runs forever, whatever it is like that ends up usually resulting in like a crash or setting off some sort of alarm. And you typically know fairly immediately that something's gone wrong. And I think the challenge with things like agents and so forth is that you can have a bad agentic loop, but it doesn't result in a crash, but it really results in just keep calling that model over and over again and trying to make adjustments and then giving you sort of plausible outputs or updates, but the whole time you're getting billed.

這和 Hugging Face 案很像:並非 AI 自行暴走,而是人做了決定。Amazon 有人把『用最多 token』設成 KPI。傳統軟體若有無限迴圈,通常會當機或觸發警報,錯誤很快就看得見;agentic loop 卻可能不斷呼叫模型、稍微調整、持續輸出看似合理的更新,完全不崩潰,而帳單一路增加。

Sean Falconer

So even outside of the wasteful like token use of maybe me using my company's token budget to do my laundry, do my grocery shopping list from next.

除了拿公司的 token 預算做洗衣、購物清單這種浪費以外……

Gregor Vand

To my laundry, that'd be.

如果 AI 真能替我洗衣服,那倒不錯。

Sean Falconer

Great. Yeah, that'd be great to do my laundry. But then there's also like legitimate excess spend where you might just end up having your agentic harness spin for some period of time where it's just churning against tokens and you don't even know that something wrong is happening. And that kind of goes back to the earlier stories as well of just, we ultimately need a lot more observability into like what is happening. Like the presumably someone at OpenAI, if they hadn't been really paying attention to like what was going on in this experiment, would have saw that this agent with the right observability tools in place is like hammering, you know, hugging face and trying the same. And that should set off certain alarms. And I think it's similar in this case where if you have like an agent loop that's out of control and spending excess tokens, like there should reasonably be some guardrails. I mean, you have that with other services, like if you spin up a particular elastic instance or something in the cloud, you're typically setting your top line provisioning of those types of things and you have some controls over it.

沒錯,真能洗衣服當然很好。但也有正當工作造成的超額支出:agentic harness 可能轉了很久,只是不停消耗 token,而你完全不知道出了問題。這又回到 observability。OpenAI 的實驗若有適當監控,就應該看得出 agent 一直猛攻 Hugging Face、重複同樣嘗試並觸發警報;失控 agent loop 也該有 guardrails 和支出上限。其他雲端服務早就會限制最高佈建量,AI 也需要類似控制。

Sean Falconer

It just feels like we haven't thought through all those controls for AI right now. And I guess part of it's just this race to try to compete everybody and everybody kind of feeling like they're behind.

目前看來,我們還沒有替 AI 想清楚這整套控制。部分原因或許是大家競爭得太急,人人都覺得自己落後。

Gregor Vand

Yeah, I'm going to just jump ahead for a second on. I wouldn't say what it is because that's a spoiler, but in one of the things I'm going to bring up on Hacker News highlights, it's a very reliable source, as you'll find out at the end. But basically this was someone who'd done a bunch of stuff with Claude, and we'll get to that. But he points out towards the end that Claude doesn't provide reliable methods of counting tokens, despite live showing token counts, reporting token consumed for sessions and billing for tokens. And he said, but I'm sure this is temporary and this will be fixed. It's just crazy that we do actually have a system at the moment where you literally just don't know what is happening and like exactly what it's going to cost and why. And as we're going to get into an open weight side of things like this is really feeding into the rise of open weight models as well.

先劇透一點我稍後的 Hacker News 焦點:有人大量使用 Claude 後指出,Claude 雖然即時顯示 token、回報 session 用量並按 token 計費,卻沒有可靠的 token 計數方式。他樂觀地說這應該只是暫時問題,但現在你真的不知道系統做了什麼、為什麼這樣收費。這也在推動開放權重模型的興起。

Sean Falconer

You think of a lot of stuff in cloud or even what we saw with ride sharing where they give you somewhat like a prediction model of what the spend will be for certain actions. So it's like, OK, well, I want to go from here to here in Uber or Lyft and I'll be like, oh, that's going to probably cost you X number of dollars. So you have some visibility in the like what the cost would be. And you can do similar things with certain cloud calculators and stuff. And it can get kind of, do we need that for AI? If I'm saying like create a engineering plan for some sort of feature, can I get an estimate of the budget required to do that and then based on what that budget is, maybe try adjusting the plan or something like that to try to optimize it down?

雲端服務或叫車平台通常會先預估支出:Uber 或 Lyft 會告訴你這趟大約多少錢,雲端也有成本計算器。AI 是否也需要?如果我要它替某功能制定工程計畫,能否先估算所需預算,再依預算調整或精簡計畫?

Gregor Vand

Yeah, that's interesting thinking sort of could you effectively put in your ask your prompt and I say prompt. I mean that almost sounds like a year ago or something. We're talking about spitting up agents etcetera. But yeah, try and get some kind of estimate before it sets off. Slightly digressing.

這很有意思:也許在請求裡先要求估價,再啟動 agent。『prompt』這個詞聽起來已經像一年前的說法了,現在談的是派出 agents。不過我們有點離題了。

4. Microsoft 的 AI 漲勢與 Apple 五兆美元時刻Microsoft's AI-Driven Surge and Apple's Trillion-Dollar Mark
Gregor Vand

So let me get us back to the headlines, which the next one we have is just the fact that Microsoft has sort of has gone back on a bit of a tear when it comes to its valuation, which is interesting. It's like one of the largest jump off shares or sorry for their fourth biggest jump on record. So it's up 16%. And we know we don't often cover just like pure financial news of tech companies, But to see Microsoft making these strides is pretty interesting. Some people might then think Oh well this is like partly to do with OpenAI, but it does own still 1/4 stake of OpenAI and claimed that that had contributed 24 billion of revenue, which was about 7% of the 332 billion in sales it reported. But a lot of it was really just AI driven revenue, unlike you've been massive investment in data centres. But actually that's been completely, in theory, vindicated by the amount of revenue they're also making.

回到新聞:Microsoft 的估值又開始狂飆,單次股價上漲 16%,是史上第四大漲幅。我們很少只談科技公司的財務,但這次相當有趣。有人會以為全靠 OpenAI;Microsoft 仍持有其約四分之一權益,並稱 OpenAI 貢獻 240 億美元收入,約占全年 3,320 億美元銷售額的 7%。更大方向是 AI 收入強勁,先前大規模投資資料中心,現在看來確實被營收證明有價值。

Sean Falconer

I think if you look at also Nadala's quotes related to the announcement of their quarter performance and so forth, and also the recent thing that we covered also in the last episode where on X he had written about how the AI companies or the model companies are charging you twice and so forth. In here, he says that every model is substitutable. He's telling essentially investors that Microsoft is deliberately building its infrastructure so it can swap out things like OpenAI for its own models or for Anthropic. And I think that it seems like they're moving towards a view where they're trying to allow essentially their customers to be very flexible and adapt, which I think is makes sense. Like I think the average enterprise now is using at least five different models, and you probably want to be able to do that so that it's a little bit like being hybrid cloud, although it's easier to be hybrid model where you get power essentially in the negotiations if you're not wholly dependent on a single vendor. And it seems like I think Microsoft's kind of leaning in that way.

從 Satya Nadella 的財報發言,以及他先前批評模型公司『向你收兩次錢』的貼文,可以看到 Microsoft 的方向。他說每個模型都可替換,等於告訴投資人:基礎設施刻意設計成能把 OpenAI 換成自家模型或 Anthropic。企業平均可能同時用至少五種模型,因此保持彈性很合理;這像 hybrid cloud,但 hybrid model 更容易。只要不完全依賴單一供應商,談判就有籌碼。Microsoft 正在朝這條路走。

Sean Falconer

But I remember back in March, Microsoft stock dipped and then everyone was like freaking out and essentially calling for the death of Microsoft. And now it's back. And I just think that the overall the market is like extremely volatile right now. There's these huge swings constantly from, I mean, IBM had its biggest drop recently, biggest single day drop in like 50 years or something like that recently. And that was coming off like a huge pop just two months earlier. So I don't know what goes up must come down. So who knows, like we might be talking about Microsoft in another quarter or two, how they had the largest single day loss in a day or something like that.

三月 Microsoft 股價下跌時,大家還在宣告它的末日,現在又回來了。市場極度波動:IBM 最近創下約五十年最大單日跌幅,兩個月前才大漲。漲久必跌,說不定一兩季後,我們又會討論 Microsoft 的史上最大單日損失。

Gregor Vand

Yeah, for sure. It is quite hard to predict as the market should be, I guess. But yeah, we're seeing just, yeah, huge swings when it comes to AI related stocks when one minute chip maker's up, one minute chip maker's down. And just another tangent there. Yeah, Apple briefly hits 5 trillion valuation, which is pretty insane. But just before recording, I I double checked and they had actually reported numbers very recently in the last I think couple of hours and they've gone back to 4.9 trillion. So I don't feel too sorry for them. But yeah, only only 4.9, but interesting to see that they're still notched above 5. And yeah, what's driving their revenue, while still they've got very good iPhone sales and they've got still very impressive services revenue and even Greater China revenue as well. But both of those services in Greater China were a little bit less than what was expected, which is again just what's kind of driven that. But it's very interesting to see that they can still, we've talked about this on previous SED news like still managed to stick on these like lines of business that are not that AI.

市場本來就難預測,但 AI 股票的擺盪尤其巨大,晶片商忽上忽下。另一個岔題:Apple 曾短暫突破五兆美元估值,錄音前公布數字後又回到 4.9 兆,倒也不用替它難過。iPhone、服務與大中華區收入依然強勁,後兩項略低於預期;有趣的是,它仍靠許多並非 AI、甚至不太沾 AI 的業務站穩。

Gregor Vand

Driven or even really like AI adjacent to be honest, very interesting, but they are having issues with memory chips, which is another sort of topic. But even in super base, we've like been told by one of our suppliers, like we have different suppliers depending on where you live to get laptops. And some of our new employees are being told like weeks before their MacBook will arrive because we do customer specs, not just like off the shelf, but yeah, now we're being told weeks, which is quite exceptional.

不過 Apple 也遇到記憶體晶片問題。就連 Supabase 的供應商也通知我們,客製規格 MacBook 可能要等好幾週,新員工拿到電腦的時間被拉長,這相當罕見。

Sean Falconer

Oh wow, yeah, I just got a a new laptop and this is my first time recording on this. So. But it's interesting, like with Apple too, a lot of their lines of business which are like incredibly successful, they're still like minority in the particular vertical. Like iPhone is wildly successful, but it's not the most dominant phone. I guess maybe from a single vendor, but Samsung might actually be bigger. But then obviously from an operating system standpoint like Android is, there's more Android devices than there are iOS devices. Similar with computers as well. Like I haven't used APC in a very long time, but it's still the dominant machine overall, I guess. But I think it's they have incredible like brand loyalty and they do make fantastic machines overall. So clearly the things that they're doing is working for them.

我剛拿到新筆電,這是第一次用它錄音。Apple 很有意思:各條業務線都非常成功,卻未必在該領域占絕對多數。iPhone 很成功,但 Android 裝置總量更多;Mac 也是,PC 仍占主流。Apple 靠驚人的品牌忠誠與優秀產品維持地位,顯然策略有效。

Gregor Vand

Yeah, absolutely.

完全同意。

5. Waymo 與 Uber 的自駕車合作走向拆夥Autonomous Vehicle Partnership Dissolves in Key US Cities
Gregor Vand

So then moving on to, we've also covered Waymo and driverless like from a few different angles. But this was interesting because we did talk about Waymo and Uber like being partners in ways and then obviously frenemies in other ways. And yeah, like, we're now actually seeing an official split of that partnership. So if you use one in San Francisco, this might sound confusing because it's a way more app, it's a way more car. So sort of where does Uber figure in that? But actually, this is for some of their other US territories. So Waymo had like first partnered with Uber apparently in May 2023, and that was to launch in Phoenix. Then it was followed by Austin and Atlanta. And that's like, so Waymo cars available through the Uber app. And then Uber managed the vehicle fleet, apparently with another partner called Avomo. But then in May, Uber and Waymo parted ways in Phoenix. And then the two companies have also clashed over the quality of Waymo services, apparently in Austin and Atlanta.

接著談 Waymo 與無人駕駛。我們以前說 Waymo 和 Uber 有時合作、有時又像亦敵亦友,現在這段合作正在正式拆夥。舊金山是 Waymo App 配 Waymo 車,看不出 Uber 在哪;其他市場則不同。Waymo 2023 年 5 月先和 Uber 在 Phoenix 合作,後來擴到 Austin 與 Atlanta:使用者從 Uber App 叫 Waymo,車隊由 Uber 與另一家夥伴管理。今年五月雙方在 Phoenix 分道揚鑣,也對 Austin、Atlanta 的 Waymo 服務品質產生衝突。

Gregor Vand

So yeah, it's kind of interesting to see they tried, but I think it was always going to be challenging to see how Waymo being owned by Google, like, how is this actually gonna now at the end, could these two actually be true partners? Or was this just always gonna be, again, a tipping point of where, like, that partnership just kind of had to end?

他們確實嘗試過,但 Waymo 畢竟屬於 Google,兩家公司能否真正成為長期夥伴一直很可疑;這段關係或許早晚都會走到必須結束的臨界點。

Sean Falconer

Yeah. I mean they say all partners are meant to be broken at some point. Like, I mean, especially where they're both in ride sharing, like clearly at some point their mutual interests are going to become too competitive to each other. Essentially in a lot of ways, I think very similar to how a lot of partnerships work where Uber was essentially supplying demand and operations while Waymo was weak on those particular spots. And then as Waymo has grown and raised more money, essentially, they don't want need those training wheels anymore. They can build out their own and build their own network and own it end to end. So I feel like this this was probably always going to be the ultimate end of that relationship.

尤其兩家都做叫車,利益早晚會變得太競爭。Uber 原本補上 Waymo 較弱的需求端和營運能力;Waymo 壯大並募到更多資金後,就不再需要這些『輔助輪』,能自己建立並端到端掌握網路。因此這很可能本來就是關係的最終結局。

Gregor Vand

Yeah. So yeah, it's sort of not officially over yet, but just people familiar with the matter apparently, again via Financial Times. But yeah, it seems pretty unsurprising really that this was going to probably break apart at some point.

目前還不算正式結束,只是 Financial Times 引述知情人士;但合作最終拆開,確實毫不意外。

6. Ray 創建者 Anyscale 加入 Neo-cloud NscaleRay Creator AnyScale Joins Neo Cloud N Scale
Gregor Vand

And then, yeah, just to wrap up on the headlines, maybe the acquisition and we were just talking before we started recording. So let me try and get this right. I believe it's any scale was acquired by N scale. And when we were talking earlier, I said, is that scale? And it's like, no, that's not scaled or AI is different again. So we've got amazing naming these days. But yeah, what? What? What's this one about?

新聞最後談一筆收購。讓我把名字說清楚:Anyscale 被 Nscale 收購;不是 Scale AI,那又是另一家公司。現在的命名實在精彩。這筆交易到底怎麼回事?

Sean Falconer

Yeah. So any scale which is known, they were the creators of Ray, which a lot of inference infrastructure and fine tuning and model infrastructure runs on a very well established open source project. Any scale was the company that tried to build or built essentially like a managed version of Ray around that and they raised like a billion dollars or so in 2022, but I think and then they just sold to N scale for 1.65 billion, which is a Neo cloud. So for those that aren't familiar with Neo cloud, Neo cloud is especially what the term is used for companies that offer primarily GPU as a service versus kind of general computing. So there's all kinds of these Neo cloud companies are now available from I'd actually talked to any scale at a variety of different times. Like, I think this was probably like a decent outcome for them because from I think it's probably hard to really grow that manage Ray as an independent company into like a really big company, but as part of a like GPU infrastructure company, it's probably a good like pairing. So I think that it makes a lot of sense. But I do like I always confuse any scale with scale AI and then the other N scale.

Anyscale 是 Ray 的創建者,Ray 是成熟的開源專案,許多推論、微調與模型基礎設施都建在上面。Anyscale 圍繞 Ray 做受管服務,2022 年估值約十億美元,現在以 16.5 億美元賣給 GPU-as-a-service 的 neo-cloud Nscale。Ray 的受管服務要獨立長成大型公司可能很難,併入 GPU 基礎設施公司反而很搭,算是不錯的結果。只是 Anyscale、Scale AI、Nscale 真的很容易混淆。

Sean Falconer

And I don't know the history of how those company names came together. But generally when the thing that people a lot of times strive for with naming companies is you want a name that you can say and people can remember. And I'm not sure they they hit the mark there where it's like N scale, scale, any scale. It's I can remember who's who in that Venn diagram.

我不知道這些名字的歷史。公司名稱理應好說、好記,但 Nscale、Scale、Anyscale 顯然沒達標;我總是記不清那張 Venn diagram 裡誰是誰。

Gregor Vand

For sure. I mean scale dot AI, OK, that's a great name to have dot AI. You can basically have anything that is your company and it's like going to sound good and it's, you know, 5 letters, but then any scale and then N scale. That's pretty confusing. But and then, yeah, just a sidebar piece of news. Scale dot AI now have a new CEO as well, which is interesting because that was founded and led by Alexander Wang, not the fashion designer, if anyone knows that one. But this is Alexander without an E at the end of the Alexander and Meta had taken us virtually just under 50% stake. That was kind of the point. So Scale and Theory are still running their own show. But when you've got I think, 49% stake from Meta, you know, you're quite beholden to them. But yeah, in that transaction, Alexander went to head up like the AI side of Meta Total. Now Scale have a new CEO. So that'll be interesting to see how that all Nets out.

Scale AI 至少拿到很棒的 .ai 網域,又短又好聽;Anyscale 和 Nscale 就真的混亂。順帶一提,Scale AI 也換了 CEO。公司由 Alexandr Wang 創辦並領導;Meta 取得略低於 50% 的股份,理論上 Scale 仍獨立,但有 49% 股權就很受 Meta 牽動。Alexandr Wang 也轉去領導 Meta 的 AI 業務,後續如何發展很值得看。

Sean Falconer

Yeah, they raised like 14 billion from of with a round led by Meta, so pretty significant.

那一輪由 Meta 領投,募得約 140 億美元,規模非常大。

Gregor Vand

Yeah. I mean actually just the new CEO of scale is actually he has a former Google cloud executive called Francis de Souza. So definitely interested to follow on with that one and see sort of how that all emits out.

Scale 新 CEO 是前 Google Cloud 高階主管 Francis deSouza;我會繼續觀察這件事最後如何發展。

Sean Falconer

Yeah, I saw also like Fireworks who just raised a huge round and made a lot of news. They their new head of engineering just came over from Google. So I think you're you're starting to see it's common pattern though. You have people who reach executive positions at large companies like Google and you know, they get may be able to board with that and then want to go back to, you know, building and moving faster.

Fireworks 最近也募到一大輪資金、新聞很多,新任工程主管來自 Google。這似乎是常見模式:有人在 Google 等大公司升到高階職位後,可能厭倦了,再回到能快速打造產品的環境。

Gregor Vand

Yeah, for sure, yeah, Fireworks, super interesting. We do have an episode with fireworks with one of the Co founders, Benny Chen. So yeah, go check that out. I think that came out around March this year, so well before this fundraising was confirmed. But yeah, they're definitely having a moment. They are, you know, sort of the Infra 4 open weight models. And that's becoming increasingly interesting to many companies for many reasons.

Fireworks 很有意思。我們大約三月訪問過共同創辦人 Benny Chen,當時融資尚未確認。他們正在成為開放權重模型的重要基礎設施,對許多公司越來越有吸引力。

7. Moonshot AI 的 Kimi 挑戰 Frontier ModelsMoonshot AI's Kimi Model Challenges Frontier AI
Gregor Vand

But yeah, that's actually quite a nice Segway into the main topic for today, which is we're kind of calling it the Kimi moment. You know, Kimi, which is an open weight model, the arch company is called Moonshot AI. So if you've heard of Moonshot AI, that's Kimi and vice versa from a Chinese Moonshot AI is a Chinese original company. They do have offices, I believe in the Valley and, and in Singapore and that kind of thing. But pretty much seen as a, a Chinese company, which sort of frames a lot of why this is quite interesting. But yeah, like when we think of open weight models, there was this the DeepSeek moment first, which we can sort of touch on. And then really in the last like almost like 3 months, we've just seen this sort of especially rapid adoption by many companies of Kimi and, and really forcing companies to sort of think differently about using, you know, foundational models from the big players or the big names rather, and getting some quite interesting and very high quality results out of Kimi. The DeepSeek moment. I think like Sean, how did that sort of kick things off when we think about open weight?

這正好帶進今天主題:『Kimi moment』。Kimi 是開放權重模型,母公司 Moonshot AI 被視為中國公司,也在矽谷、新加坡等地設點。談開放權重時,先有 DeepSeek moment;近三個月 Kimi 又快速被企業採用,迫使大家重新思考是否一定要用那些最知名的基礎模型。它提供了很高品質的結果。Sean,DeepSeek 當初如何替開放權重揭開序幕?

Sean Falconer

Yeah. I mean, I think the DeepSeek moment ended up having sort of more like market impact in some sense, because I think prior to that, the sense was that you couldn't really get these really powerful models except from these like Frontier labs like Anthropic and Google OpenAI and so on. And so it was kind of shocking when DeepSeek came out and the Nvidia's market cap like cratered as a result of that. And of course it's, you know, certainly that was back since then, but and then when you look at the Kimi launch, which which is the largest open weight model ever released, you're beating some of the top close models on particular benchmarks. It seemed like nobody really panicked. And I think part of that is not because it's less impressive, but it's because we've kind of gotten used to the idea that these like Chinese lab overweight models are competitive and it's so it's less of a shock essentially we've been desensitized to it.

DeepSeek 帶來較大的市場震撼。在那之前,大家以為只有 Anthropic、Google、OpenAI 等 frontier labs 才做得出強大模型,所以 DeepSeek 出現時甚至讓 NVIDIA 市值重挫。之後 Kimi 發表史上最大的開放權重模型,部分 benchmark 勝過頂尖閉源模型,市場卻沒恐慌;不是它不驚人,而是大家已習慣中國實驗室的開放權重模型具有競爭力,震撼感已降低。

Sean Falconer · 需對照原音

But I do think that what the market reaction has been more around, and I think this is something we're already coming to is like where this was somewhat on the back of what happened with like with the Methos and Fable And the reaction to that where people got scared that they could invest in a model and then have potentially like a foreign government say like you can't use this model anymore. Then I think that has increased the interest in diversification of models and then also having an open weight model strategy along with having a sort of closed weight model strategy.

市場真正的反應比較像模型被禁用的風險。企業若投資某模型,卻可能被政府突然禁止使用,就會更重視模型多元化,並在閉源模型策略之外,同時建立開放權重策略。

Gregor Vand

Yeah. And we'll probably touch on this sort of the banning of models because of course it comes into this as well, if you want to say ironically, because, yeah, when you ban at least pause banner, I'm a foundational model, a closed source model if you like, then open weight becomes very interesting. But then a lot of open weight comes from China. So we're going to kind of be interested in that. Why sort of the last couple of months has Kimi really just like exploded, I think in interest and popularity, it's kind of moved from, oh, it's good enough for professionals to this kind of actually competes with GPT and Clods, you know, and it's kind of closed that gap in basically weeks, at least from what has been sort of released.

我們也會談禁用模型的問題。若某個閉源基礎模型遭暫停或封禁,開放權重就很吸引人;偏偏大量開放權重模型又來自中國。Kimi 在短短幾個月從『專業工作勉強夠用』變成真正能和 GPT、Claude 競爭,幾週內迅速縮小差距,這正是值得研究的地方。

Sean Falconer

Yeah, and they also released like a ton of models on like back-to-back, essentially. They're moving super, super fast.

而且他們接連發表大量模型,速度非常快。

Gregor Vand

Yeah, exactly. So I mean, yeah, if if you sort of look at the actual lineage, I guess like they've been doing sort of quarterly drops if you like, you know, July 2025 was K2 and then within 12 months you've gone through 2.52.62.7 and now K3. That's a credibly rapid cadence. Like if you look at those jumps as being as significant as jumps that you might see from any of the foundational players, but definitely not, but not on that cadence effectively.

看發布脈絡,他們大約按季更新:2025 年 7 月是 K2,十二個月內一路經過 2.5、2.6、2.7 到 K3。若每次躍進都接近大型基礎模型廠的版本升級,這個 cadence 快得驚人。

Sean Falconer

Yeah, Yeah. I mean, I think the model distillation, the practice has really speed up how quickly people are coming up with models and and models that are comparable in performance, which was a big part of of the conversation around though like initial DeepSeek launch. I do think that is part of the some of the kind of criticism on some of the benchmark results that we've seen from Kimi is that in particular, like there was a lot of headlines around their MCP tool calling performance, but that beating Opus for example. But in particular, that benchmark is relatively new and there is, I don't know if this is more just sort of like jealousy in the dialogue or whether there's some truth to this. But you can essentially sort of bias towards performing really well on certain benchmarks, whether it's MCP one or it's, you know, the MLLU style benchmarks as well. And then you can have like really good benchmark performance, but it might not actually match like reality.

Model distillation 確實加速了模型推出與追平效能,這也是 DeepSeek 初次發布時的重要話題。不過 Kimi 的 benchmark 也遭批評,例如 MCP tool calling 成績勝過 Opus;這個 benchmark 很新,可能被針對性最佳化。我不確定批評是嫉妒還是真有根據,但無論 MCP 或 MMLU,都可能出現 benchmark 很漂亮、實務表現卻不相符的情況。

Sean Falconer

That's where some of the criticism has been on even some of the tests of, you know, testing a model against LSTAT and stuff like that and then saying, oh, it's it's, you know, outperformed lawyers on the LSAT, But then like the Lsat's actually not necessarily a good indication of what a lawyer does on a day-to-day basis and stuff. So these are some of the the nuance and the some of this is like general, you know, model criticism. But this is some of the dialogue that I've heard around Kimi in particular is like, are they building to optimize for the metric? Kind of like what we talked about in the Amazon story is like, you know, whatever the KPI is, people are going to try to optimize for that KPI. So you always have to be careful essentially what the metric is that you're measuring people against.

同樣地,模型在 LSAT 上勝過律師,不代表它能做好律師日常工作。這是一般模型評測都要注意的細節,也正是 Kimi 討論的一部分:它是否只是針對指標最佳化?就像 Amazon 的 KPI 故事,人總會追逐被衡量的數字,所以必須慎選 metric。

Gregor Vand

Yeah. And on June 12th, K 2.7 code was released, you know, and this is obviously very much the software engineering tuned model. And as you were saying, yeah, that became kind of the benchmark moment, especially through MCP. And, and that was a correct tool invocation is sort of how that's how that's measured apparently, like scoring in theory, you know, 81% and that was versus say Opus 4.8 at 76%, which is a pretty, if you believe in these benchmarks, that's a pretty meaningful jump. But then the distribution piece was kind of interesting, like GitHub actually made 2.7 generally available by Copilot, but there was kind of like an interesting piece there where for enterprise teams on Copilot business and enterprise, like it was actually this model was off by default and admins had to explicitly enable it. And Github's change log flag this and sort of said it may be less aligned than other Copilot models, which is interesting because obviously they have a their own interests being, you know, part of the Microsoft OpenAI ecosystem. But, you know, obviously they couldn't miss not providing that to users given that it's on Azure and they they still get inference from it.

6 月 12 日發布的 Kimi K2.7 Code 是針對軟體工程調校的模型,也在 MCP 正確工具呼叫 benchmark 上成為焦點:理論分數 81%,高於 Opus 4.8 的 76%。若相信 benchmark,差距很有意義。GitHub 很快讓 K2.7 進入 Copilot,但 Business 與 Enterprise 預設關閉,管理員須明確啟用;changelog 還警告它『可能比其他 Copilot 模型更不 aligned』。GitHub 身處 Microsoft–OpenAI 生態系,措辭背後顯然也有自身利益;但模型已在 Azure 上,不能完全不提供。

Gregor Vand

But yeah, very interesting that this sort of slightly odd warning came with it. That wasn't really clear what that was about.

這個略顯奇怪、又沒有清楚解釋的警告很值得注意。

Sean Falconer

Yeah, I, I believe it's the first Chinese lab open weight model that's been built into to copilot. I do get like having, you know, working with a lot of enterprise customers, they do really careful about which models they use. So I can kind of understand having some level of governance there. But of course, there's a certain biased interest from Microsoft point of view of how do you, you know, craft the language around this particular model. But there is, I think that fear with the enterprise. So I understand some level of controls there, but obviously there's also a certain bias that can be been put into it. I think one of the things that's interesting about this is Moon Shot as well. As you know, this goes for all the Chinese labs.

我想它是第一個進入 Copilot 的中國實驗室開放權重模型。企業客戶確實很謹慎選模型,所以某種治理控制可以理解;但 Microsoft 如何描述競爭模型也難免有偏見。企業的顧慮是真實的,控制合理,只是其中也可能摻入立場。另一點是 Moonshot 和所有中國實驗室共同面對的限制。

8. 晶片禁運如何逼出中國實驗室的創新How Chip Embargoes Push Chinese Labs to Innovate
Sean Falconer

They don't have these like top tier NVIDIA chips available to them like the US labs use because of the various chip embargoes.

因為晶片禁運,他們無法像美國實驗室一樣取得最頂級的 NVIDIA 晶片。

Gregor Vand

We don't think they do, but yeah.

至少我們認為他們拿不到。

Sean Falconer

Even if they did have some, probably not at the scale. So there's a certain scarcity of resources that they have to deal with. And I think that in some ways like that is forcing them to innovate in a way that maybe the team, the US based companies or the Western based companies where they don't have that scarcity of resource aren't necessarily forced to do that. So some ways like maybe we're creating the thing that we fear. But good news of that is it's it's probably forcing the other companies, like the western based companies to react to this, to lower token prices and maybe also think about how to they keep costs down. It's a little bit, you know, if you look at Google's beginnings, Google started as a research project at Stanford and you have students essentially didn't have a lot of resources. So they'd be very, very creative about how they scaled Google. And then even that carried over to when Google did initially have some funding.

即使拿得到,規模大概也不一樣。資源稀缺迫使他們創新,而資源充足的美國或西方公司未必承受同樣壓力;某種程度上,我們可能親手創造了自己害怕的競爭者。好消息是,這也逼西方公司降低 token 價格、思考成本。Google 最初只是 Stanford 的研究計畫,學生資源有限,所以不得不創造性地擴展系統;即使後來拿到資金,這種文化仍延續。

Sean Falconer

But if that had a project had started out of like a bigger, more established, well funded company, they probably wouldn't have been ripping apart like cheap machines and like wiring them using Legos and like stuffing them together as quickly as in to compact the size within the data center as much as possible. And they would have just bought like, you know, super beefy servers. And the downside of that would have been like a lot of the innovation that we've seen since then in the in cloud infrastructure kind of came from some of those like forcing yourself to deal with these scarcity resources. And my hope with all this is we see a similar thing of innovation driven from the scarcity resources that carry over to the Frontier Labs in other parts of the world as.

若 Google 一開始就在資金雄厚的大公司裡,他們可能只會購買昂貴伺服器,不會拆便宜機器、用 Lego 固定、想辦法把設備塞進資料中心。正是資源限制,催生後來許多雲端基礎設施創新。我希望今天由稀缺推動的創新,也能擴散到世界其他 frontier labs。

Gregor Vand · 需對照原音

Well, yeah, I mean, it's it's interesting sort of just sidebar into like she was even sort of behind the the founder Yang Zillen. He's sort of probably was known as Yang the genius. There was a very good profile on him in the Financial Times actually last weekend. He's 34 years old. He was known at university for being like just as much into music and having academic sort of soirees, if you like. And actually Kimi, when he first released it, you know, as a chat bot, it it didn't do very well. It was sort of outages and this kind of thing. And so it's clearly as you're calling out, it's like managed to catch up despite not having the same access or we don't think couldn't possibly have the same kind of access despite what it still may have. Back to sort of then the turning point here like July 16th, which is for sort of history that's then 34 days after 2.7 it dropped K3 drops and that was a model with 2.8 trillion parameters, which is I believe the largest open weight model released to date. And I think what's interesting here is then on the July 27th, the full model weights were actually published.

Moonshot 創辦人楊植麟也很有意思,Financial Times 最近替這位 34 歲、被稱為『楊天才』的人做了專訪。他在大學時熱愛音樂,也辦學術沙龍。Kimi 最早作為聊天機器人推出時表現並不好,還頻繁中斷;但即使沒有同等晶片資源,它仍追了上來。7 月 16 日,K2.7 發布僅 34 天後,K3 問世,擁有 2.8 兆參數,據說是歷來最大的開放權重模型;7 月 27 日又公開完整權重。

Gregor Vand

And this sort of then leads into this whole like provenance piece, which is when companies are starting to be asked like, well, you know, you're using AI to generate so much information within your own companies or analyze datics, you want to actually know how this was determined and so on and so forth. And here we are. We actually have a competitive model with the closed foundational models with the full weights published on Hugging Face, which is like massively impressive to me. That's like a huge turning point when when you've actually got like effectively all the weights right there. And this isn't just, you know, a model that takes up some tiny little specific task. It's it's really competing now.

這帶出 model provenance。公司用 AI 生成或分析大量內部資訊時,開始被問:結果如何形成?現在我們有一個能與閉源基礎模型競爭、完整權重又公開在 Hugging Face 的模型,這是巨大轉折。它不是只會做單一小任務,而是真正具備全面競爭力。

Sean Falconer

I think overall this is whether Moonshot and Kimi kind of win out or not. I think overall this is like good for consumers of these models because it will force the other companies to essentially react to this. And hopefully actually I saw a headline today that OpenAI was reducing some of their token costs by up to 80%. So and we saw similar thing after DeepSeek get drove down token costs as well. And then all the stuff we were talking about earlier of like token maxing and companies kind of getting sensitive to how much they're spending. Like overall, I think it'll be good for consumers of these things. You know, one thing I was thinking about with this story too, you know, we've covered a lot of Meta over the last year and their decisions in terms of like hiring and what they're trying to do with their AI lab.

不論最後是否由 Moonshot 與 Kimi 勝出,消費者都會受益,因為其他公司必須回應。OpenAI 據說正把部分 token 價格最多降低 80%,DeepSeek 之後也發生過類似降價。企業正對 token-maxxing 和支出變敏感,競爭整體上是好事。這也讓我想到 Meta:我們過去一年談了很多它的招聘與 AI 實驗室策略。

9. 模型多元化與 Meta 的 Llama 路線Diversification of Models and Meta's Llama Approach
Sean Falconer

But Meta had such a head start in this like open weight model space with the llama models and it's been a long time since I've heard anything about the llama models.

Meta 明明靠 Llama 在開放權重領域起步很早,但我已經很久沒聽到 Llama 的消息。

Gregor Vand

Funny you said that because, yeah, there was something, I didn't dig into it, but I just saw the headline, which was that Zuckerberg is basically lobbying to ensure that Chinese models don't get banned. And to me, that only meant one thing, which is like, well, they have to keep that door open because like, Llama's not maybe going where he hoped it was.

巧的是,我剛看到 Zuckerberg 正遊說不要封禁中國模型。我的直覺是:他必須讓這扇門保持開啟,因為 Llama 的進展可能不如預期。

Sean Falconer

I think like his instinct of trying to own the open weight model space was probably right. The execution was bad. I don't know what happened internally there, but like they were clearly like on a good path, especially with what we're seeing now between the Chinese labs, between like what's happening with fireworks and the other inference providers that have focused on open weight. Like there's clearly like a huge Tam available for open weight to own a big part of the business. And I think realistically, especially in the enterprise, most enterprise businesses will probably have. A mixture of both open weight investments that they've done and then as well as the closed source models. And they'll probably be the strategy that many, many companies take for some, you know, period into the future. I don't know how you know Meta ended up and maybe they'll bounce back, but you know, they were on to something, but they you know, they haven't been able to execute.

他想掌握開放權重市場的直覺大概是對的,執行卻出了問題。中國實驗室、Fireworks 與其他 open-weight inference providers 都證明這是一個巨大市場;尤其企業很可能長期同時投資開放權重與閉源模型。Meta 原本走在正確道路上,卻沒有執行成功;也許仍有機會反彈。

Gregor Vand · 需對照原音

I think that's really good. You know, analysis, it is the right strategy, wrong time. Like, you know, too early effectively perhaps or just like too early in the given they would join idiot from the US be asked to kind of produce too much too soon and like that can distort things. So who knows, maybe like a rug or Mesa with llama they they'll kind of adopt more of a Kimi approach to make this succeed. Who knows. But yeah, I mean, just to kind of wrap up on Kimi and sort of this open weight obviously resurgence, but like, yeah, maybe search model provenance is a real thing that's being sort of talked about now when it comes to, you know, it's a business risk, right? Like you're putting money and basing your company on whether it's like for coding or for other tasks. But you are now having to really decide like it's a procurement question on what are you buying effectively? And I'm like, what are the risks with that? Like could it get too expensive or could you move it on to your own infra if you really needed to or not your own infra, but exactly via fire, you could still own that sort of end to end there.

這個分析很好:策略正確,但時機或內部節奏錯了,也可能是太早、承受過多美國方面的壓力,導致方向扭曲。也許 Meta 會讓 Llama 重整,改採更像 Kimi 的方式。Kimi 與開放權重復興最後帶回 model provenance:企業把資金與產品押在模型上,這就是採購與風險問題。你買的是什麼?會不會漲得太貴?必要時能否移到自己掌握的 infra,或透過 Fireworks 端到端控制?

Gregor Vand

That's interesting pricing, we've talked about that like pricing is just kind of getting out of control. And it was really a case of, like, when, not if, is this going to, like, stop being possible for, you know, anthropic and OpenAI to kind of, like, charge this way because it's just not sustainable for many businesses. And, yeah, and like, the speed as well, the speed of iteration is like, clearly being a huge piece here where, however Moonshot are doing it, they're just iterating like crazy.

定價已失控;Anthropic 與 OpenAI 不可能永遠這樣收費,因為許多企業承受不了。迭代速度也很關鍵:不管 Moonshot 怎麼做到,他們正以瘋狂速度更新。

10. JetBrains、重構、CodePen 與 DoomQLJetBrains, Refactoring, CodePen, Doom QL, and Future AI Trends
Gregor Vand

So on to our favorite parts, Hacker News highlights. Do you want to go first, Sean?

接著進入我們最喜歡的 Hacker News highlights。Sean,你先來?

Sean Falconer

Sure, so this headline really caught my eye and made me made me laugh. Which was the speaking agents like Cavemen save 65% of tokens we test. So this was by Jetbrains and they benchmarked the caveman skill, which is a Claude Code skill that makes it respond and terse caveman speak to save tokens and they claim 65% token savings. So they essentially put that to the test. It was focused on agentic coding tasks and they found actually in reality it was about 8.5% output token savings versus 65%. But a big part of that is they were focused on coding tasks where a lot of it's going to be code. You can't turn that into caveman speak. You got diffs, you got tool calls, so the caveman skill might not be best for that versus, you know, some sort of more conversational chat. But I just really loved it. Reminded me of like some ignoble prize, you know, research out there. There. It's just kind of a ridiculous topic.

有則標題讓我大笑:『讓 agents 用穴居人語氣說話可省 65% token——實測』。JetBrains 測試一個 Claude Code skill,用極短的穴居人式語句節省 token;宣稱可省 65%,但在 agentic coding 任務裡,實際輸出 token 只少 8.5%。原因是大部分輸出是程式碼、diff 和 tool calls,無法改成穴居人語。它也許更適合一般對話,但這個題目本身就像搞笑諾貝爾獎研究一樣荒謬可愛。

Gregor Vand

That's like, yeah, super funny. It's interesting because I didn't actually scan ahead. So that is actually a little bit similar to what I picked out, which is called the economic benefits of refactoring. So this is on the fairly popular martinfowler.com website posted by Java E user on Hacker News. So thanks for that. But this is not written by Martin Fowler himself. But basically this was, you know, someone else that thought works and it was, can you basically decrease especially the input token cost if you refactor your code or like what are the consequences of that? And I mean, the TLDR is yes, if you refactor then basically you can dramatically reduce your input tokens. And kind of the theory behind this was that the saving is because the agent has to read less code. But the bit that might not be like that sounds obvious, but it's actually not because there is less code to read. It's actually that the overall code in a certain layer they used to help test this, like it stayed constant. But the fact that the agent is then able to successfully identify smaller subsets of files to read is like the key bit there.

真巧,我選的主題很接近:『重構的經濟效益』,文章登在 martinfowler.com,由 Thoughtworks 的作者撰寫。問題是:重構能否降低 input token 成本?答案是能,而且幅度很大。不是因為總程式碼量變少;測試中某一層的程式碼總量其實不變,而是 agent 重構後能更準確找到需要閱讀的較小檔案集合。

Gregor Vand

It's like refactoring into smaller chunks, but also making sure that those chunks are very clearly DRY, don't repeat yourself, etcetera, etcetera. So that the code that the agent needs to go and grab is smaller. You can't just sort of chunk it up into small chunks and then go, well, it's smaller. Like it then doesn't even know it has to still get all the small chunks because it doesn't know what's most important. But the refactoring is what then turns it into having that context of having like what is most important, go and find that small chunk, input that small chunk. And the output token and output code was virtually the same in this experiment. So which is super interesting. There's a small sort of side note on that, which was just him saying that Claude, as Claude is actually not good at refactoring. And I can definitely attest to that. But if you can go through the motions of get it to refactor, then this is quite could be quite a huge saving for anyone who's trying to like reduce their input token costs.

重點是把程式碼拆成小而清楚、符合 DRY 的單元,讓 agent 知道哪一塊最重要,只讀那一小部分。隨便切成碎片沒有用,它仍可能得把所有小塊都讀完。實驗中的 output token 與產生的程式碼幾乎不變,節省的是 input。作者也提到 Claude 並不擅長重構,我可以作證;但若能完成重構,降低 input token 成本的效果可能很大。

Sean Falconer

Yeah. So basically, bottom line, good design leads to also optimize token cost.

所以結論是:良好設計也會最佳化 token 成本。

Gregor Vand

Funny that. Yeah, we just go back to how we used to design software and.

真有意思,我們又繞回傳統上設計軟體的方法。

Sean Falconer

I mean, it's the same like if you think about the human cost, like if you have like poorly designed software, then there's going to be more human cost each time someone unfamiliar with it needs to like ramp up and make some sort of change.

人力成本也是同一回事。軟體設計很差時,每個不熟悉系統的人要上手並修改,都會付出更多時間。

Gregor Vand

Exactly.

完全正確。

Sean Falconer

And then you have a another one on CodePen.

你還選了一則 CodePen 的新聞。

Gregor Vand

My second one was yeah, just thanks to user Robin Riyala posting the fact that code Pen 2.0 has come out. This takes me back. That's I was quite interested in this. I don't know if anyone else out there, this makes it sound terrible. So if you like CodePen, this is sorry about this. I just, I haven't been doing a lot of front end work for a long time, but used to love CodePen, used to, you know, put up all sorts of things on there and a really useful tool as well. Like I did a little coding like class. I went back to my old high school like a long time ago and did like a coding class. And Codepen was amazing because you could just get people spun up in a browser writing front end code and just see it do its thing straight there. They didn't have to build with like we didn't have to get them spun up with some sort of repo or anything. So it was really helpful. But yeah, I mean, CodePen must have come out like probably 15 years ago or something like that. So CodePen 2 point O I'll just quickly run through a couple of things that they call out that they have files and folders now. So it's interesting.

對,感謝 Hacker News 使用者 Robin Rendle 分享 CodePen 2.0。這讓我很懷舊;我很久沒做大量前端工作,但以前很愛 CodePen。我曾回母校開短期程式課,用它讓學生直接在瀏覽器寫前端、立即看到結果,不必先架 repository,極其方便。CodePen 大概已推出十五年,如今 2.0 終於加入檔案與資料夾。

Gregor Vand

It's almost becoming a bit IDE esque, but files and folders and then they have sort of they do, you know, build steps anyway, but they've now like added this concept of blocks, which looks quite nice where you can sort of see exactly which bits are in this build process or like the linting process. So that's kind of fun yeah, real time collaboration as well. So finally, multiplayer on CodePen. So for anyone that uses CodePen a lot, I'm sure that's quite a huge uplift. So all right, you've come and checked out. Yeah. Still a fun place to go and just experiment with front end D stuff. So yeah. And then just a special one, I guess not technically through Hacker News, but thanks to Ilya Roshetnikov for tweeting at me. And Sean, we do often cover these Doom. Can you run Doom on something? And thanks to Ilya, he pointed out that there's another one called Doom QL, which is basically using an SQL query is the frame buffer, which is like, I think that definitely rivals, you know, dim on TypeScript types. Yeah, I think that's probably the the closest I can think of that it rivals. So thank you, Ilya, for shouting that one out to us. That was very, very fun to read through.

它越來越像 IDE,也加入 blocks 來呈現 build、lint 等流程,還有即時協作,多人終於能一起用 CodePen。它依舊是試驗前端的好地方。最後一則沒上 Hacker News:Ilya Roshchepkin 告訴我 DoomQL,用 SQL query 當 framebuffer 來跑 Doom。這足以媲美用 TypeScript types 跑 Doom;謝謝 Ilya 分享,讀起來非常有趣。

Gregor Vand

So yeah, that's it for another SED News. Have we got any looking ahead predictions, Sean, which we usually get wrong?

又一集 SED News 到此結束。Sean,有沒有照例可能猜錯的未來預測?

Sean Falconer

Yeah. I mean, I think the safe predictions here would be that we're going to see more headlines on the token maxing issue as companies start to adjust. I think we'll see more also conversations around open weight versus closed model, diversification of models. I think those are both going to be big topics conversation through to the end of the year.

保守預測是:企業開始調整後,token-maxxing 會有更多新聞;開放權重與閉源模型、模型多元化也會持續成為年底前的重要話題。

Gregor Vand

Yeah, for sure. I will then say, well, because we just talked about iteration and Kimi, let's assume that by this time next month, we're already on like Kimi 3.2 or if I'll just push the boat out, give me 3.53 point 5 by end of August. Let's see if that lands. So thanks everyone for for tuning in as always. And we'll be back next month with our SED News.

同意。既然談到 Kimi 的迭代速度,我大膽猜下個月此時已經有 Kimi 3.2,甚至八月底就到 3.5。看看會不會命中。謝謝大家收聽,我們下個月再帶來 SED News。

Sean Falconer

Thanks everyone, cheers.

謝謝大家,再見。

建議練習方式

  1. 先讀章節導讀與詞彙,建立內容地圖。
  2. 隱藏中文,播放一個章節並閱讀英文。
  3. 第二次播放時顯示中文,確認沒聽懂的內容。
  4. 隔天不看文字重聽,口頭回答理解題。

官方節目頁 · Spotify 單集