Reflecting on One Year at Bedrock: Transformative LLM re:InventingBedrock 一周年回顾:变革性的 LLM re:Invent 之旅

Hits

"The power of a machine learning model is not just in its size, but in how well it can generalize as it scales."「机器学习模型的力量不仅在于其规模,更在于它随规模扩展时的泛化能力。」Geoffrey Hinton, recipient of the Nobel Prize in Physics (2024); A.M. Turing Award (2018).Geoffrey Hinton,2024 年诺贝尔物理学奖得主;2018 年图灵奖得主。




re:Invent wrapped up, and this week marks the one-year milestone for me at Bedrock. As I reflect on the past year, it's a good moment to look back at both the highlights and the hurdles, the moments of happiness and those of frustration. Notably, Bedrock celebrated its first anniversary in October, a significant milestone that highlights the team’s rapid growth and impact. Having been involved in building and founding several products at Amazon, I feel fortunate to be part of Bedrock, a primitive service and cutting-edge product, and to grow alongside it. It’s been an exciting journey of growth, collaboration, and learning, and I’m grateful for the opportunity to contribute to such a transformative initiative. Challenging and fulfilling.


Anthropic Collaboration: Building Something Bigger Together

As one of the first two engineers when the project funded to build serverless inference for Anthropic models on Bedrock. One of the most significant aspects of my time at Bedrock has been working closely with our external partner, Anthropic, the leading frontier model provider company worldwide. From the start, the team I worked at under Bedrock inference has been the only team working directly with Anthropic and this unique partnership has involved a lot of collaboration work. I remember an interesting moment when a L7 colleague slack-messaged me saying, "Andy, good to meet you. I’ve heard your name mentioned even by [XXX] (the co-founder of Anthropic), when we were working on this ... project." It’s always humbling to hear that your name is being discussed at such high levels.

Our collaboration with Anthropic has been deep and ongoing. Anthropic is one of the companies I have collaborated with; operating with zero overhead and has a rare and appealing atmosphere of rising. This partnership led to some truly amazing product launches, such as the latest intelligent models, Computer Use, Haiku 3.5 Speed SKU on Trainium 2, and last week's "The Claude in Amazon Bedrock Course". I vividly recall several late nights where both teams worked tirelessly to debug issues, even staying up the entire night. Of course, the happy hours afterward at Seattle’s SLU for a successful launch were just as memorable.


Privilege, Confidentiality, and Exclusivity

Working so closely with Anthropic has come with a significant level of responsibility and confidentiality. My skip-level manager's teams report directly to VP in the reporting chain. Everything we do is considered privileged and confidential, can't be shared beyond my skip-level manager's teams, even within the same department, and we have secured codebases, documents, channels, basically everything related to our work. This sense of exclusivity and responsibility reminds me of my time in Alexa (Device org), where access to certain departments required specific clearance.

For nearly a year, our team’s work has been shrouded in secrecy. We’ve been unable to share much with others outside our direct teams, including other Bedrock teams. The good news is that people recognize the work we’ve done in delivering Claude models, and while the details are closely guarded, the results speak for themselves. Over time, we’ve become more accustomed to operating in this high-trust, high-privacy environment. And you'll find a dedicated pair of L10s (Distinguished Engineers), L8s (Senior Principals), and L7s (Principals) ICs participating in team operational meetings as well as partner and customer calls, executive titles like these are rarely seen in other product teams.


Claude on Trainium

Thanks to teams effort. This has been a significant milestone, and I feel privileged to have been deeply involved in several key aspects of this project over the past year. There were many unknowns and challenges along the way, but through careful setup, testing, and benchmarking, we worked closely with the Anthropic and Annapurna teams to achieve great results. The outcome turned out to be a big win, and it was exciting to see the announcement spread during re:Invent, especially after Peter’s "Monday Night Live" session. What’s interesting is a coincidence that occurred during this project: Because of my long-standing part-time contributions to the Linux Foundation community (not work relevant), I was invited to visit San Francisco as a member of the PyTorchCon Program Committee. While there, I participated in a panel discussion where James Bradbury, Head of Compute at Anthropic, was also a speaker. The following week, back in Seattle, I worked on tasks that happened to involve several interactions with James on the Claude on Trainium. It felt like a serendipitous moment when you know someone outside of work and later find yourselves working together full-time on the same project.


Milestones and Deliverables: A Year of Innovation

(re:Invent 2024 Amazon Bedrock launches)


I'm at Bedrock inference team, a team behind this year's most trending keywords, including 'Claude', 'Trainium' and 'Computer Use'. Looking at the work we’ve accomplished, the list of deliverables is nothing short of impressive. In the past year, we’ve launched several net new innovations and features that have pushed the boundaries of AI and model inference. Some of the key milestones include:

  • December: Amazon Bedrock launches with Claude 3.5 Sonnet in the AWS Top Secret cloud
  • December: Amazon Bedrock announces preview of prompt caching. (re:Invent)
  • December: Introducing latency-optimized inference for foundation models in Amazon Bedrock. (re:Invent)
  • November: Anthropic and Palantir Partner to Bring Claude AI Models to AWS for U.S. Government Intelligence and Defense Operations.
  • November: Claude 3.5 Haiku (GA) on Amazon Bedrock.
  • October: Upgraded Claude 3.5 Sonnet model from Anthropic, introducing groundbreaking Computer Use capabilities.
  • September: Claude models Provisioned Throughput.
  • August: Amazon Bedrock offers select FMs for batch inference at 50% of on-demand inference price.
  • July: Claude 3.5 Sonnet launched.
  • June: Amazon Bedrock now available in London, São Paulo, and Canada (Central) regions.
  • May: Claude 3 Sonnet and Haiku launched in Frankfurt.
  • April: Anthropic’s Claude 3 Opus model now available on Amazon Bedrock.

  • I’ve had the privilege of contributing to or leading these launches, and the pace of innovation here is intense. Each month brings one or two major launches, and this is just the external-facing work. Behind the scenes, there are countless internal deliverables that contribute to the overall momentum. While quarterly or yearly deliveries may have been the norm in the past, at Bedrock, it’s all about the monthly or bi-weekly cadence, with a constant push to meet what some may call "unreasonable" timelines. You can never know what is the character of the next model. The complexity and effort required to meet these fast-paced expectations should never be underestimated.


    Move Fast

    Bedrock and the broader GenAI department are known for their fast-paced environment. One thing I’ve observed as a new joiner is that people, rarely reply to Slack messages. Initially, this felt uncomfortable, and I wondered if it was a sign of something personal. However, as I settled in, I realized that it’s simply a matter of time. There's just not enough bandwidth to keep up with every message, even when you want to support your team.

    In some ways, I've come to appreciate the return-to-office (RTO) setup, especially now that we have newly assigned desks in our Nitro North office building. With people in the office, it's easier to find colleagues and have quick, direct conversations. Junior team members can easily approach senior colleagues to ask questions and get help to overcome obstacles. This has improved our communication efficiency significantly. Another interesting observation is the prevalence of cold calls in our culture. It’s not unusual to receive calls at any time from your colleagues, skip-level managers, or even GMs. On the very first day I joined Bedrock, I was already hosting design reviews and being asked to jump into tasks. This unique style of working can be intense but also fosters a sense of urgency and immediate engagement.

    An AI researcher shared his typical daily schedule at OpenAI in a tweet. It's interesting to see how others in the same field organize their day. Balancing work and life is always a challenge, especially in a high-pressure environment like Bedrock. But I find that the excitement of working on such a transformative product makes it all worthwhile. The product we’re building has the potential to shape industries and change the way we interact with technology, and that gives me a deep sense of fulfillment.


    Team 3x

    My manager's team 3x, with more people expected to join in the coming months. Bedrock is expanding rapidly, and with that growth comes the challenge of managing increasing complexity, new team members, and cross-country collaborations. As our team grows, balancing workloads and maintaining effective communication becomes a delicate art. Peer manager partnerships and collaboration are crucial, especially when coordinating across different time zones and cultural contexts.

    Bedrock is actively expanding its team, and I’ve had the chance to conduct many interviews. At Amazon, I’ve participated in over 110 interviews, with 40 of them specifically for Bedrock, averaging about one interview per week over the year.


    Building Primitive Services

    Peter DeSantis gave a highly comprehensive session at re:Invent this year, where he provided an in-depth explanation of the complexities involved in machine learning workloads, covering both model training and inference. I won't dive into the specifics of prefill and token generation. For model inference, getting a model up and running is no easy task. Bedrock plays a heavy lifting role in facilitating collaboration between compute resources, enabling access to the most advanced and intelligent models for utilization. I’ve enjoyed working with the foundational-level building blocks like LLM (Large Language Models) inference. In 2023 letter to shareholders, Andy Jassy explained and shared the concept of "Primitive Services". My previous experience spans various AWS products such as S3, Containers, and Virtulization, and I’ve found that working on primitive services offers a unique sense of satisfaction. Someone says Bedrock is the "Lambda of LLM". I say it is more than that. Mart Garman stated in an interview with SiliconANGLE, "Inference is the next core building block. If you think about inference as part of every application, it becomes integral, just like databases." It’s exciting to see how the generative AI space is evolving, and I’m thrilled to be part of Bedrock at such a transformative moment in AI history.


    What Comes Next

    As we continue to grow, I foresee a future where AI models become even more deeply integrated into our daily lives. At NeurIPS, Ilya Sutskever, remarked that "Pre-training as we know will end." So, what comes next ? "Agents", "Synthetic data", and "Inference time compute ~O(1)". FM / LLM of today may increasingly resemble the everyday software applications we use, with future updates and iterations becoming as routine as downloading a new browser version from the App Store. In this highly competitive landscape, the version numbers of models from different providers could eventually resemble browser version numbers, reflecting regular updates and improvements. We also expect the boundary of knowledge within these models to shrink, making them accessible to a broader audience. Jason Clinton, CISO Anthropic, shared at re:Invent, more innovations are expected to emerge in AI agents. Anthropic published new research this week "Building effective agents". "2025 will be the year of agentic systems".

    Anthropic has done a significant amount of impactful work here, particularly with recent Model Context Protocol (MCP) and its Computer Use (apologies for mentioning it repeatedly, but it's an incredibly powerful feature). One of the papers I read about its potential exploration “The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use” by Show Lab at National University of Singapore, is definitely worth checking out. These efforts have played a crucial role in initializing the entire ecosystem.


    AI Safety

    Responsible AI and AI safety are unavoidable topics, spanning both technology and sociology domains. Both Anthropic and Amazon have made serious investments in this area. It is intriguing to read the AI safety research findings done by Chatterbox Labs. Working closely with Anthropic has been particularly insightful, as their deep understanding and dedication to AI trust and safety are truly impressive. This concentration is evident in the interview with Anthropic’s CEO Dario Amodei on the Lex Fridman Podcast and their renowned article "Anthropic's Responsible Scaling Policy," which introduces the concept of "AI Safety Levels (ASL)." These efforts underscore their commitment to ensuring the safety of artificial intelligence.

    Additionally, in a Financial Times interview, Dario highlighted two key priorities for the company in 2025, one of which is "mechanistic interpretability, looking inside the models to open the black box and understand what’s inside them.." He described this as "perhaps the most societally important".

    I couldn’t agree more: interpretable machine learning and AI safety are complementary and mutually reinforcing. Adopting a cautious yet optimistic stance toward current advancements in AI is both wise and necessary, and I look forward to continued progress in this field.


    Overall, my first year at Bedrock has been an incredible journey. I’ve learned so much, faced new challenges, and contributed to innovative products that are making a real impact. I’m excited to see what the future holds as we continue to grow, innovate, and push the boundaries of what’s possible in the world of AI and cloud computing.


    A heartfelt thanks to Ben, Tim, Tommy, Evan, Nova, Jin, Frank, Josh, Peter, Justin, and all the Anthropic teams we collaborated with, and everyone at Bedrock, my managers, Ravi, Shashank, and all my amazing colleagues. It has been an absolute privilege to collaborate and work alongside this year.

    re:Invent 落下帷幕,而这一周也正好是我加入 Bedrock 一周年的日子。回顾过去这一年,此刻正适合盘点那些高光与坎坷、快乐与沮丧的瞬间。值得一提的是,Bedrock 在十月迎来了它的一周岁生日——这个重要的里程碑印证了团队的快速成长与影响力。 在 Amazon 参与构建和创立过多个产品之后,我很庆幸能加入 Bedrock 这样一个原语级服务(primitive service)和前沿产品,并与它共同成长。这是一段充满成长、协作与学习的精彩旅程,我很感激能为如此具有变革意义的项目贡献力量。充满挑战,也充满成就感。


    与 Anthropic 的合作:共同构建更宏大的事业

    当项目立项、要为 Bedrock 上的 Anthropic 模型构建 Serverless 推理时,我是最早的两名工程师之一。在 Bedrock 这一年里,最重要的经历之一就是与我们的外部合作伙伴 Anthropic——全球领先的前沿模型公司——紧密合作。从一开始,我所在的 Bedrock 推理团队就是唯一直接与 Anthropic 对接的团队,这种独特的伙伴关系带来了大量的协作工作。我记得一个有趣的时刻:一位 L7 同事在 Slack 上给我发消息说:"Andy,很高兴认识你。我们做这个……项目的时候,连 [XXX](Anthropic 联合创始人)都提到过你的名字。"听到自己的名字在如此高的层面被讨论,总是让人既惶恐又欣慰。

    我们与 Anthropic 的合作深入且持续。Anthropic 是我合作过的公司中运转几乎零开销的一家,有着一种少见而迷人的上升气息。这一伙伴关系促成了一些真正令人惊叹的产品发布,比如最新的智能模型、Computer Use、基于 Trainium 2 的 Haiku 3.5 Speed SKU,以及上周的 "The Claude in Amazon Bedrock Course" 课程。我清楚地记得有几个深夜,两边团队为了排查问题通宵达旦地一起工作。当然,成功发布后在西雅图 SLU 的 happy hour 同样令人难忘。


    特权、保密与专属

    与 Anthropic 如此紧密的合作,也意味着相当高的责任与保密要求。我隔级经理(skip-level manager)的团队在汇报链上直接向 VP 汇报。我们做的一切都属于 privileged and confidential(特权与机密),不能分享到隔级经理团队之外——即使是同一部门内部也不行。我们的代码库、文档、频道,基本上所有与工作相关的东西都是加密隔离的。这种专属感和责任感让我想起在 Alexa(设备组织)的日子,那时访问某些部门也需要特定的许可。

    近一年来,我们团队的工作一直笼罩在保密之中。我们无法与直属团队之外的人分享太多,包括其他 Bedrock 团队。好消息是,大家都认可我们在交付 Claude 模型方面所做的工作——虽然细节被严格保密,但成果本身就是最好的证明。随着时间推移,我们也越来越适应这种高信任、高隐私的工作环境。你会发现有专门的 L10(Distinguished Engineer)、L8(Senior Principal)和 L7(Principal)IC 成对地参与团队运营会议以及合作伙伴和客户电话——这样级别的头衔在其他产品团队中是很少见的。


    Claude on Trainium

    这要归功于团队的共同努力。这是一个重要的里程碑,我很荣幸在过去一年里深度参与了这个项目的多个关键环节。一路上有许多未知和挑战,但通过细致的搭建、测试和基准评测,我们与 Anthropic 和 Annapurna 团队紧密合作,取得了出色的成果。最终结果是一场大胜,尤其是在 Peter"Monday Night Live" 环节之后,看到这一发布在 re:Invent 期间广泛传播,令人兴奋。有趣的是,这个项目期间还发生了一个巧合:由于我长期以业余身份为 Linux Foundation 社区做贡献(与工作无关),我受邀前往旧金山,担任 PyTorchCon 程序委员会成员。在那里,我参加了一场小组讨论,Anthropic 的计算负责人(Head of Compute)James Bradbury 也是讲者之一。而接下来那一周,回到西雅图后,我手头的工作恰好涉及与 James 就 Claude on Trainium 的多次交流。这种感觉很奇妙——你在工作之外认识了某个人,之后却发现你们全职共事于同一个项目。


    里程碑与交付:创新的一年

    (re:Invent 2024 Amazon Bedrock 发布一览)


    我所在的 Bedrock 推理团队,正是今年最热门关键词——'Claude'、'Trainium' 和 'Computer Use'——背后的团队。回顾我们完成的工作,交付清单令人印象深刻。过去一年里,我们发布了多项全新的创新和功能,不断突破 AI 与模型推理的边界。其中一些关键里程碑包括:

  • 十二月:Amazon Bedrock 携 Claude 3.5 Sonnet 登陆 AWS Top Secret 云
  • 十二月:Amazon Bedrock 宣布 prompt caching 预览版。(re:Invent)
  • 十二月:Amazon Bedrock 推出面向基础模型的延迟优化推理。(re:Invent)
  • 十一月:Anthropic 与 Palantir 合作,将 Claude AI 模型引入 AWS,服务于美国政府情报与国防业务。
  • 十一月:Claude 3.5 Haiku 在 Amazon Bedrock 正式发布(GA)。
  • 十月:Anthropic 升级版 Claude 3.5 Sonnet 模型发布,引入突破性的 Computer Use 能力。
  • 九月:Claude 模型 Provisioned Throughput。
  • 八月:Amazon Bedrock 为部分基础模型提供批量推理,价格为按需推理的 50%。
  • 七月:Claude 3.5 Sonnet 发布。
  • 六月:Amazon Bedrock 上线伦敦、圣保罗和加拿大(中部)区域。
  • 五月:Claude 3 Sonnet 和 Haiku 在法兰克福上线。
  • 四月:Anthropic 的 Claude 3 Opus 模型登陆 Amazon Bedrock。

  • 我有幸参与或主导了这些发布,这里的创新节奏非常紧凑。每个月都有一两个重大发布,而这还只是对外可见的工作。幕后还有无数内部交付在为整体势能添砖加瓦。过去按季度或按年交付或许是常态,但在 Bedrock,一切都以月度甚至双周为节奏,持续冲刺那些有人称之为"不合理"的时间线。你永远无法预知下一个模型的特性是什么。达成这种快节奏预期所需的复杂度和投入,永远不应被低估。


    快速前进

    Bedrock 乃至整个 GenAI 部门都以快节奏著称。作为新人,我观察到一个现象:大家很少回复 Slack 消息。起初这让我不太舒服,甚至怀疑是不是针对我个人。但随着逐渐适应,我意识到这纯粹是时间问题——即使你想支持你的团队,也确实没有足够的带宽去跟上每一条消息。

    从某种程度上说,我开始欣赏重返办公室(RTO)的安排,尤其是现在我们在 Nitro North 办公楼有了新分配的工位。大家都在办公室时,更容易找到同事进行快速、直接的交流。初级团队成员可以很方便地向资深同事请教问题、扫清障碍。这显著提升了我们的沟通效率。 另一个有趣的观察是我们文化中普遍存在的"cold call"。同事、隔级经理、甚至 GM 随时可能给你打电话,这并不稀奇。加入 Bedrock 的第一天,我就已经在主持设计评审、被要求直接上手任务了。这种独特的工作方式强度很大,但也营造出一种紧迫感和即时投入的氛围。

    一位 AI 研究员在一条推文中分享了他在 OpenAI 的典型日程。看看同领域的其他人如何安排一天,很有意思。平衡工作与生活始终是个挑战,尤其是在 Bedrock 这样高压的环境中。但我发现,参与构建如此具有变革性的产品所带来的兴奋感,让这一切都值得。我们正在打造的产品有潜力重塑行业、改变人与技术的交互方式,这给了我深深的成就感。


    团队 3 倍扩张

    我经理的团队扩大了 3 倍,未来几个月还会有更多人加入。Bedrock 正在快速扩张,随之而来的是管理日益增长的复杂度、新成员融入以及跨国协作的挑战。随着团队壮大,平衡工作负载和保持有效沟通变成一门精细的艺术。同级经理之间的伙伴关系与协作至关重要,尤其是在跨时区、跨文化背景协调的时候。

    Bedrock 正在积极扩充团队,我也因此参与了大量面试。在 Amazon,我累计参加了超过 110 场面试,其中 40 场专门为 Bedrock 进行——平均下来这一年每周约一场。


    构建原语服务

    Peter DeSantis 在今年 re:Invent 上做了一场内容极其全面的分享,深入讲解了机器学习工作负载的复杂性,涵盖模型训练与推理。这里我就不展开 prefill 和 token 生成的细节了。就模型推理而言,让一个模型跑起来绝非易事。Bedrock 在协调计算资源方面承担着繁重的底层工作,让最先进、最智能的模型得以被使用。我很享受与 LLM(大语言模型)推理这类基础性构建模块打交道。在 2023 年致股东信中,Andy Jassy 阐释并分享了"原语服务(Primitive Services)"的理念。我之前的经历横跨 S3、容器、虚拟化等多个 AWS 产品,我发现做原语服务有一种独特的满足感。有人说 Bedrock 是 "LLM 界的 Lambda"。我认为它远不止于此。Matt Garman 在接受 SiliconANGLE 采访时表示:"推理是下一个核心构建模块。如果你把推理看作每个应用的一部分,它就会变得不可或缺,就像数据库一样。"生成式 AI 领域的演进令人振奋,我很高兴能在 AI 历史上如此关键的变革时刻身处 Bedrock。


    接下来会是什么

    随着我们不断成长,我预见 AI 模型将更深地融入我们的日常生活。在 NeurIPS 上,Ilya Sutskever 表示"我们所熟知的预训练将会终结"。那么接下来是什么?"Agents"、"合成数据",以及"推理时计算 ~O(1)"。今天的基础模型 / LLM 可能会越来越像我们日常使用的软件应用——未来的更新迭代会像从 App Store 下载新版浏览器一样稀松平常。在这个高度竞争的格局中,不同提供商的模型版本号最终可能会像浏览器版本号一样,反映着持续的更新与改进。我们也预期这些模型的知识边界会不断收敛,让更广泛的人群能够使用。Anthropic 的 CISO Jason Clinton 在 re:Invent 上分享说,AI Agent 领域将涌现更多创新。Anthropic 本周发布了新研究 "Building effective agents"。"2025 将是 Agent 系统之年"。

    Anthropic 在这方面做了大量有影响力的工作,尤其是最近的 Model Context Protocol(MCP)和 Computer Use(抱歉反复提到它,但它确实是一个极其强大的功能)。我读过的一篇探讨其潜力的论文——新加坡国立大学 Show Lab 的《The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use》——非常值得一读。这些工作在整个生态的启动初始化中发挥了关键作用。


    AI 安全

    负责任的 AI 与 AI 安全是绕不开的话题,横跨技术与社会学两个领域。Anthropic 和 Amazon 都在这一领域进行了认真的投入。阅读 Chatterbox Labs 完成的 AI 安全研究结果很有启发。与 Anthropic 的紧密合作尤其让我受益——他们对 AI 信任与安全的深刻理解和专注令人印象深刻。这种专注在 Anthropic CEO Dario Amodei 做客 Lex Fridman 播客的访谈,以及他们著名的文章《Anthropic's Responsible Scaling Policy》(提出了"AI 安全等级(ASL)"概念)中都可见一斑。这些努力彰显了他们对确保人工智能安全的承诺。

    此外,在《金融时报》的一次采访中,Dario 强调了公司 2025 年的两大优先事项,其中之一是"机制可解释性(mechanistic interpretability)——深入模型内部,打开黑箱,理解其中的运作机制"。他称这"或许是对社会最重要的事情"。

    我对此深表认同:可解释的机器学习与 AI 安全是相辅相成、互相促进的。对当前 AI 的进展保持谨慎而乐观的态度,既明智也必要,我期待这一领域的持续进步。


    总的来说,在 Bedrock 的第一年是一段不可思议的旅程。我学到了很多,迎接了新的挑战,并为正在产生真实影响的创新产品贡献了力量。我很期待看到未来的样子——我们将继续成长、创新,不断拓展 AI 与云计算世界的可能性边界。


    衷心感谢 BenTimTommyEvanNovaJinFrankJoshPeterJustin,以及所有与我们合作过的 Anthropic 团队,还有 Bedrock 的每一位伙伴、我的经理们、RaviShashank,以及所有出色的同事。今年能与大家协作共事,是我莫大的荣幸。