Career · June 2026职业生涯 · 2026年6月

ML Job Interviews: The Ultimate Guide机器学习求职面试:终极指南

I thought it might be helpful to write about my experience finding a job as a Research Scientist after a PhD in Machine Learning. There's almost zero information out there on this, and I wish someone had written it when I was starting out. I hope this is useful whether you're in the thick of it or just thinking about getting started.我想分享一下我在机器学习博士毕业后找研究科学家工作的经历,可能会对大家有所帮助。关于这方面的信息几乎为零,我真希望自己刚开始时有人写过这样的内容。希望无论你正处于求职高峰期,还是仅仅在考虑开始,这篇文章都能对你有用。

My process was, overall, successful: I received offers from every company I completed interviews with including: DeepMind (which I accepted), Isomorphic Labs, Cohere, Meta, and a startup in stealth. A few caveats to the first claim: Anthropic, Mistral, and TeslaAI got back to me too late and I didn't complete those processes. ReflectionAI, the one genuine rejection: they didn't like me for the RS role but switched me to their Engineering track instead.总体而言,我的求职过程很成功:我完成了面试的每一家公司都给了我录用通知,包括:DeepMind(我接受了)、Isomorphic Labs、Cohere、Meta 以及一家未公开的初创公司。但需要说明几点:Anthropic、Mistral 和 TeslaAI 回复我太晚,我没能完成这些流程。ReflectionAI 是唯一真正拒绝我的:他们认为我不适合研究科学家的职位,但让我转到了工程方向。

Most companies I applied to invited me for interviews, with the exception of: SpaceXAI, Waymo and Wayve. For SpaceXAI, I did let a friend write my application as a joke, but I didn't think it was that bad. For Waymo and Wayve: I love self-driving cars. I applied to Waymo every six months throughout my PhD (internships, then full-time) and never heard back once, despite people in my own lab getting replies. Waymo, if you're reading this, I'm open to forgiveness. My love cover letters were works of art. You can reach me at the email you already have on file, multiple times.
我想分享一下我在机器学习博士毕业后找研究科学家工作的经历,可能会对大家有所帮助。关于这方面的信息几乎为零,我真希望自己刚开始时有人写过这样的内容。希望无论你正处于求职高峰期,还是仅仅在考虑开始,这篇文章都能对你有用。 总体而言,我的求职过程很成功:我完成了面试的每一家公司都给了我录用通知,包括:DeepMind(我接受了)、Isomorphic Labs、Cohere、Meta 以及一家未公开的初创公司。但需要说明几点:Anthropic、Mistral 和 TeslaAI 回复我太晚,我没能完成这些流程。ReflectionAI 是唯一真正拒绝我的:他们认为我不适合研究科学家的职位,但让我转到了工程方向。 我申请的大部分公司都邀请了我面试,除了 SpaceXAI、Waymo 和 Wayve。对于 SpaceXAI,我确实让一个朋友帮我写了申请作为玩笑,但我没想到结果会这么糟。至于 Waymo 和 Wayve:我非常喜欢自动驾驶汽车。读博期间我每六个月申请一次 Waymo(先是实习,然后是全职),但从未收到过回复,尽管我实验室里有人收到了回复。Waymo,如果你们看到这篇文章,我愿意原谅你们。我写的求职信堪称艺术品。你们可以用存档的邮箱联系我,多次。

Getting Interviews获得面试机会

Getting interviews is its own challenge, and if you're struggling with it, the levers are the usual ones: more papers, trendier topics, and better internships. You can look at my CV on my website for reference, but briefly: I had 4 first-author (or co-first) papers from my PhD, published at ICLR / NeurIPS / ICML, covering a mix of trendy topics (LLMs, RL) and less fashionable ones (Meta-Learning, Evolution Strategies). I also had an internship at Apple and prior industry experience as a Software Engineer at Meta. If I had to give a rough benchmark: 3+ first-author papers and at least one internship or industry role seems to be the threshold for consistently getting callbacks at top labs.获得面试机会本身就是一项挑战。如果你在这方面遇到困难,通常的发力点包括:发表更多论文、研究更热门的课题、以及获得更好的实习经历。你可以参考我网站上的简历,简单来说:博士期间我发表了4篇第一作者(或共同第一)论文,发表在ICLR/NeurIPS/ICML上,涵盖了热门话题(LLMs、RL)和不太时髦的方向(元学习、进化策略)。此外,我还在苹果公司实习过,之前在Meta做过软件工程师。如果非要给出一个粗略的标准:3篇以上第一作者论文加上至少一次实习或行业经验,似乎是持续获得顶级实验室面试机会的门槛。

That said, if you're already getting interviews: more papers will not help you at this point. You need to pass the interviews, and often the people interviewing you won't even look at your CV. So, stop focusing on your research and your papers, and start focusing on interview prep! I understand the feeling of wanting to postpone, but you are never going to feel ready, so just start prepping now.不过,如果你已经获得了面试机会:此时再发表更多论文也无济于事。你需要通过面试,而面试你的人通常甚至不会看你的简历。所以,不要再把精力放在研究和论文上了,开始专注于面试准备吧!我理解你想要推迟的心情,但你永远不可能觉得准备好了,所以现在就立刻开始准备。

A few other details worth knowing: cover letters, referrals, cold emails, and LinkedIn/X.其他值得了解的细节:求职信、推荐、冷邮件和LinkedIn/X。

LinkedIn / X: A lot of companies advertise roles here, and for internships in particular it's sometimes the only way to apply. You have to fill out a Google form linked from the post for your application to actually count. Follow the people you admire at the companies you're interested in so you don't miss these.LinkedIn / X:很多公司会在这里发布职位信息,尤其是实习岗位,有时这是唯一的申请途径。你需要填写帖子中链接的谷歌表单,申请才算有效。关注你感兴趣的公司中你敬佩的人,这样就不会错过这些信息。

Referrals: Nice to have, but not necessary. At DeepMind I had a referral for two roles and none for a third, I got invited to interview for one referred role and the unreferred one. At Anthropic, I heard nothing until I discovered an ex-collegue of mine had recently joined and asked him to put in a referral for me. So, worth getting if you can, but don't let the absence of one stop you from applying.推荐:有则更好,但不是必须的。在DeepMind,我有两个职位的推荐信,第三个没有,结果我收到了一个推荐职位和那个无推荐职位的面试邀请。在Anthropic,我一直没有回音,直到发现一位前同事最近加入了该公司,并请他为我推荐。所以,如果有推荐机会值得争取,但不要因为没有推荐就放弃申请。

Cold emails: Emailing the hiring manager or someone on the team directly (if you know who they are) is often appreciated. Don't just repeat your CV (you can attach it), use the email to explain why you'd be a good fit for that specific team and what genuinely excites you about their work. For this: at Deepmind I emailed my Hiring Manager, he was happy about it and replied. For another role, I saw the Hiring Manager explicitly encourage people to email him on X… but I could just not be bothered, since I was already going through interviews with the other team. I ended up getting an interview with them despite not sending the email.冷邮件:直接给招聘经理或团队中的某人发邮件(如果你知道他们是谁)通常会受到欢迎。不要只是重复你的简历(可以附上),要用邮件解释你为什么适合这个特定团队,以及他们的工作究竟哪里让你兴奋。例如:在DeepMind,我给招聘经理发了邮件,他很高兴并回复了。对于另一个职位,我看到招聘经理在X上明确鼓励大家给他发邮件……但我当时懒得发,因为我正在另一个团队面试。结果我没发邮件也获得了面试机会。

Cover letters: Rarely required, but worth doing properly when they are. Please, for the love of everything I hold dear, do not just ask Claude / Gemini / ChatGPT to write it for you. You can absolutely write it yourself and then ask one of them to polish it, that's fine. But try to make some of your personality and excitement shine through.求职信:很少要求,但如果有要求就值得认真对待。拜托,千万不要直接让Claude/Gemini/ChatGPT帮你写。你可以自己写,然后让它们帮忙润色,这没问题。但一定要让你的个性和热情闪耀出来。


Companies: Startups vs Big Tech公司:初创公司 vs 大科技

The summary is: it depends, more than any generic pros/cons list I can write will tell you. But here are the main factors to consider.总结是:这要视情况而定,比我所能列出的任何通用优缺点清单都要复杂。但以下是主要考虑因素。

Finding startups is harder. There's no central place to look. Ask your labmates, friends, and former colleagues... Word of mouth is the best way to find good ones in your area of research. At the same time, because they are harder to find, competition is generally less fierce for these positions.找到初创公司更难。没有一个集中的地方可以查找。向你的实验室伙伴、朋友和前同事打听……口碑是找到你研究领域内好初创公司的最佳途径。同时,因为难以找到,这些职位的竞争通常不那么激烈。

Interview processes vary more at startups. Big tech follows a fairly predictable structure, while startups do their own thing. The difficulty level is comparable on average, but variance is high. Pay attention to what the interview process tells you about the company: if it feels too easy, that might be a signal about the complexity of the work you'd actually be doing. As with the interviews, I think there's simply more variance in the quality of work and people at startups compared to big tech.初创公司的面试流程差异更大。大科技公司的流程相当可预测,而初创公司则各有各的做法。平均难度相当,但方差很大。注意面试流程能告诉你关于公司的哪些信息:如果感觉太简单,那可能暗示了你实际工作内容的复杂性。与面试一样,我认为初创公司的工作和人员质量方差也比大科技公司大。

The work itself can go either way. At the right startup, the research might be more interesting and more impactful than anything you'd work on at a big lab. But it can also come with more pressure, more engineering and infrastructure work you didn't sign up for, and a research agenda that shifts often. Ask questions in interviews! Who decides research priorities? What is the path to profitability? Who are the competitors? What if OpenAI wakes up tomorrow and decides to do the same thing?工作本身则各有优劣。在合适的初创公司,研究可能比你在大型实验室做的更有趣、更有影响力。但也可能带来更大的压力、更多你没有预料到的工程和基础设施工作,以及频繁变化的研究方向。在面试中要多问问题!谁决定研究优先级?盈利路径是什么?竞争对手是谁?如果OpenAI明天醒来决定做同样的事情怎么办?

Room for growth. Startups generally offer more opportunity to grow quickly, take on responsibility, and shape the direction of the work. At a big lab you're one of many, at a startup you're much more visible. This can be a big deal (I know it was for me).成长空间。初创公司通常提供更多快速成长、承担职责和塑造工作方向的机会。在大型实验室你只是众多人中的一员,而在初创公司你更引人注目。这可能很重要(对我来说确实如此)。

CV matters. OpenAI or Anthropic on your CV is immediately recognised by anyone. A stealth startup nobody has heard of requires explanation. That's not a reason to avoid startups (and if you're a founder type, then it's completely different) but it's worth factoring in, honestly.简历价值。OpenAI或Anthropic出现在你的简历上,任何人都能立刻认出。而一个没人听说过的未公开初创公司则需要解释。这不是回避初创公司的理由(如果你是创始人类型,那就完全不同),但诚实地讲,这值得考虑。

On "security": I'm not going to make that argument. Big tech has done mass layoffs without blinking many times over the past years, neither path is 100% safe.关于“安全感”:我不打算争论这个。大科技公司在过去几年里多次大规模裁员,两者都不是100%安全的。

A note on compensation: RSUs vs stock options关于薪酬的说明:RSUs 与股票期权

This took me an embarrassingly long time to understand, so I'll try to save you the confusion. I'm speaking to UK law and taxation here.我花了很长时间才理解这一点,所以我想帮你们省去困惑。这里我指的是英国法律和税收。

With RSUs (typical at big tech), you receive actual shares in the company on a vesting schedule. When they vest, you can sell them or hold them. About half get sold immediately to cover income tax (because yes, to the absolute shock of 23-year-old me getting her first Meta shares... RSUs count as income).对于RSUs(通常在大科技公司),你会按照归属计划获得实际的公司股份。股份归属后,你可以出售或持有。大约一半会立即卖出以支付所得税(因为,是的,让我23岁第一次获得Meta股份时绝对震惊的是……RSUs算作收入)。

With stock options (typical at startups), you're not getting shares. You're getting the opportunity to buy shares at a fixed price X, regardless of what the market price Y is at the time. If Y > X, great! You can exercise your option, buy at X, sell at Y, and pocket the difference. If Y < X, your options are worthless.对于股票期权(通常在初创公司),你获得的是以固定价格X购买股份的机会,而不管当时的市场价格Y是多少。如果Y > X,太好了!你可以行使期权,以X买入,以Y卖出,赚取差价。如果Y < X,你的期权就一文不值。

Here's where it gets wild. Stock options typically expire 90 days after you leave the company. If the company isn't publicly traded yet, you can't sell your shares after buying them, meaning you might have to spend (X × options) in cash to exercise (=use) them, with no guarantee you'll ever be able to sell. And in the UK, the moment you exercise your options, you owe income tax on the Y−X difference even if you haven't sold a single share and haven't seen a penny of that money yet.接下来就更疯狂了。股票期权通常在你离开公司后90天到期。如果公司尚未上市,你买入股份后无法出售,这意味着你可能需要支付(X × 期权数量)的现金来行使期权,且无法保证将来能卖出。而在英国,你行使期权的那一刻,即使没有卖出一股、还没有看到一分钱,你也要为Y−X的差额缴纳所得税。

So if you leave a startup after two years of working there, exercise your options, and the company isn't yet public, you would have to pay for: the cost to buy the shares (X × options) plus income tax on the paper gain ((Y−X) × options × your tax rate). Before you've made anything.所以,如果你在一家初创公司工作两年后离开,行使期权,而公司尚未上市,你将需要支付:买入股份的成本(X × 期权数量)加上账面收益的所得税((Y−X) × 期权数量 × 你的税率)。在你还没赚到钱之前。

A few caveats: most companies offer a cashless exercise option, where you hand back some options to cover the cost of exercising the rest. Many also have liquidity events where they buy back some stock. But remember: each new funding round dilutes your shares, and any gains beyond the income-tax-liable portion get hit with capital gains tax (~20%) on top. Liquidity events typically value your stock below the official company valuation too.几点说明:大多数公司提供无现金行使选项,你可以交回部分期权来支付行使剩余期权的成本。许多公司还有流动性事件,会回购部分股票。但请记住:每一轮新的融资都会稀释你的股份,而超出所得税部分的任何收益还要缴纳资本利得税(约20%)。流动性事件通常以低于公司官方估值的价格评估你的股票。

Summary:总结: when a recruiter quotes you a total compensation number that includes startup equity, smile politely and mentally discount it significantly. Most likely scenario: you are not retiring next year. But I wish you the best of luck.总结:当招聘人员向你报出一个包含初创公司股权的总薪酬数字时,礼貌地微笑并在心里大幅打折。最可能的情况是:你明年不会退休。但我祝你一切顺利。

Interview Structure面试结构

Most companies follow roughly the same structure, though the weight given to each stage varies a lot.大多数公司的结构大致相同,但每个阶段的权重差异很大。

Recruiter screen. This is usually just a chill, low stakes chat. It's an opportunity to show your skills are relevant to the job, and that you actually know and can talk about the papers you are an author of招聘人员筛选。这通常只是一次轻松的、低风险的聊天。这是一个展示你的技能与职位相关、并且你确实了解并能谈论你所著论文的机会。

Technical interviews. This is the bulk of the process and where preparation matters the most. Expect 3-8 of these interviews depending on the company:技术面试。这是流程的主体,也是准备工作最重要的部分。根据公司不同,预计有3-8轮这样的面试:

  • Coding. LeetCode-style problems, typically Medium or Hard difficulty编程。LeetCode风格的问题,通常是中等或困难难度。
  • ML coding and debugging. Implementing attention, writing backward passes, spotting bugs in training loops机器学习编程与调试。实现注意力机制、编写反向传播、发现训练循环中的错误。
  • ML knowledge. Fundamentals, theory, applied ML, system design机器学习知识。基础、理论、应用ML、系统设计。

Behavioural interviews. They split into two flavours, classic behavioural questions (“tell me about a time you had a conflict”, "tell me a time you received feedback"...), research-style interviews ("what topics are you interested in?", "where do you see the field going?"). These are more casual compared to the technical interviews, but make sure to not underestimate these interviews and reflect on / prepare your answers. 行为面试。分为两种类型:经典的行为问题(“请讲述一次你遇到冲突的经历”、“讲述一次你接受反馈的经历”……)、研究型面试(“你对哪些话题感兴趣?”“你认为这个领域将如何发展?”)。这些比技术面试更随意,但切勿低估,需要认真思考并准备你的回答。


How I Prepared: Technical我的准备方法:技术部分

This is the key part! Don't skip this.这是关键部分!不要跳过。 I know extremely impressive researchers who were rejected in interviews simply because they didn't prepare. Working with ML day in and day out is not the same as being ready to implement attention from scratch, derive the backward pass, or code flash attention. Allocate at least至少 a month of regular study time.这是关键部分!不要跳过。我认识一些非常出色的研究人员,仅仅因为准备不足而在面试中被拒。日复一日地做机器学习工作,与准备好从头实现注意力机制、推导反向传播或编写闪存注意力算法是完全不同的。至少分配一个月的固定学习时间。

One meta-strategy that helped me: I did little generic prep. Almost everything was targeted at the next specific interview or company. This kept me focused and meant the material l was asked about was fresh on my mind (essential for a goldfish like me). By the end, you'll have covered most of the material anyway. But I know people that prefer to adopt different strategies, so do whatever works for you!一个对我有帮助的元策略:我很少做通用的准备。几乎所有准备都针对下一次具体的面试或公司。这让我保持专注,并确保被问到的问题在我脑海中记忆犹新(这对记性差的我至关重要)。到最后,你自然也会覆盖大部分内容。但我知道有人更喜欢不同的策略,所以选择适合你的方式!

By the end of my interview journey I'd built up a pretty comprehensive directory of resources and strategies, which I'll share below. I know, it's a lot. But the reality of ML Research Scientist / Engineer interviews is that you can be asked almost anything, from basic concepts like overfitting, to LeetCode, to implementing a transformer from scratch, to specific questions about fairly modern architectures (Griffin, TransformerXL, S4). The list reflects that range.面试旅程结束时,我已经建立了一个相当全面的资源和策略目录,我将在下面分享。我知道内容很多。但ML研究科学家/工程师面试的现实是,你几乎可能被问到任何问题,从过拟合等基本概念,到LeetCode,到从头实现Transformer,到关于相当现代架构(Griffin、TransformerXL、S4)的具体问题。这份清单反映了这个范围。

Flashcards闪卡

For ML fundamentals, applied ML, and research discussions. I tried Anki first and didn't get on with it, physical flashcards worked much better for me. More importantly, writing your own cards is half the learning, don't just download someone else's deck. I'll link my full topic list at the end if you want a starting point. When reviewing, be curious: ask yourself questions and make sure you deeply understand each topic. Multiple times while studying, I asked myself questions I was later asked in interviews. Better to solve doubts beforehand!用于机器学习基础、应用和研究讨论。我先试了Anki,但不太适应,实体闪卡对我效果好得多。更重要的是,自己写卡片就是学习的一半,不要只下载别人的牌组。我会在最后附上我的完整主题列表,如果你想以此为起点。复习时,要好奇:问自己问题,确保深刻理解每个主题。在学习的多次过程中,我问了自己后来在面试中被问到的问题。最好提前解决疑惑!

LLM mock interviews (Claude / Gemini)LLM模拟面试(Claude / Gemini)

Before each interview, I'd paste a description of the role, interview and company into my favourite LLM for the day (usually Claude) and ask it to interview me. There was surprisingly frequent overlap between those practice questions and what interviewers actually asked. I'd recommend doing this for every interview. If the difficulty feels off, start a new chat and specify your level and background more explicitly. I think Claude was the best LLM for learning, and I thought its feedback was generally fair, Gemini was a bit too flattering (“they are lucky to be interviewing you” ~ cit Gemini).每次面试前,我会把职位描述、面试和公司信息粘贴到我当天最喜欢的LLM(通常是Claude)中,让它面试我。这些练习题与面试官实际提出的问题之间有着惊人的重叠。我建议每次面试都这样做。如果难度不合适,开启新的对话,更明确地说明你的水平和背景。我认为Claude是最好的学习LLM,它的反馈总体上很公平,Gemini有点过于讨好了(“他们能面试你是他们的幸运”——引用自Gemini)。

LeetCode / NeetCodeLeetCode / NeetCode

Do at least Blind 75 and optionally NeetCode 150, focusing on Mediums. Try to get the optimal solution for each question, an O(N²) solution for TwoSum doesn’t count as a solution. Don't spend much time on Hards. If you haven't done LeetCode before: it will feel awful at first, you'll feel stupid, that's normal and it passes. By question 100 you'll feel much more confident. Just make sure to know the basic patterns (DFS, BFS, Graphs, Backtracking, DP, Binary Search...) and make sure you can implement them confidently and very quickly. Target 20 minutes or less per Medium. If you're stuck for more than 15 minutes, look up the solution, understand it, flag it for review, and move on. Breadth matters more than depth here! I did around 150 Mediums total至少完成Blind 75,如果可能再做NeetCode 150,重点放在中等难度上。争取每道题都找到最优解,TwoSum的O(N²)解法不算解。不要在困难题上花太多时间。如果你以前没做过LeetCode:一开始会感觉很糟糕,你会觉得自己很笨,这很正常,而且会过去。做到第100题时你会自信得多。只要确保掌握基本模式(DFS、BFS、图、回溯、动态规划、二分搜索……),并且能够自信、非常快速地实现它们。目标:每道中等题20分钟或更少。如果卡住超过15分钟,就查答案,理解它,标记为待复习,然后继续。广度比深度更重要!我总共做了大约150道中等题。

Books书籍

  • Designing Machine Learning Systems by Chip Huyen: covers a lot of the fundamentals and applied ML questions you'll encounter. Highlight and take notes.《设计机器学习系统》——刘知远(Chip Huyen)著:涵盖了你会遇到的许多基础和应用机器学习问题。阅读时高亮并做笔记。
  • The JAX Scaling Book: I found this after my interviews, unfortunately, but it's excellent and I'd have used it heavily.《JAX扩展策略手册》:我是在面试之后才发现的,可惜了,但它非常出色,如果早点看到我会大量使用。
  • Reinforcement Learning by Sutton & Barto: only if you're new to RL. I think it's overkill if you're already working in the area.《强化学习》——萨顿与巴托:仅当你对强化学习不熟悉时阅读。如果你已经在从事该领域,我认为这本书有些过度了。

Courses课程

  • Linear algebra: Gilbert Strang's lectures on YouTube. You can get through the whole course at 2x speed in under a day (I may be speaking from experience). He's the only reason I passed linear algebra in my undergrad. May he live another 91 years.线性代数:Gilbert Strang在YouTube上的讲座。你可以在不到一天内用2倍速学完整门课程(可能我是经验之谈)。他是我大学时通过线性代数的唯一原因。愿他再活91年。
  • Diffusion / Flow Matching: the MIT and Stanford courses are both good but quite math-heavy. If you're not actively researching in this area, questions will likely be superficial, so just get a high-level intuition and memorise the basics (e.g. diffusion SDEs and flow matching ODEs).扩散模型/流匹配:MIT和斯坦福的课程都不错,但数学内容较多。如果你不是积极研究这个领域,问题可能很浅显,所以只需要高层次直觉并记住基础知识(例如扩散SDE和流匹配ODE)。

ML coding and debuggingML编程与调试

This is where I found the fewest good resources, and where actual experience matters most. The debugging interviews were especially hard to practice, LLMs couldn't reliably generate convincingly buggy code when asked. Reviewing your own codebase (or a friend's) is probably your best bet. DeepML has some good questions but there is also a lot of useless ones. I also found these Tensor Puzzles quite helpful. The baseline you should aim for:这是我发现资源最少的部分,也是实际经验最关键的部分。调试面试尤其难以练习,LLM无法可靠地生成看起来可信的错误代码。回顾你自己的代码库(或朋友的)可能是最好的方法。DeepML有一些好问题,但也有很多无用的。我也发现这些Tensor Puzzles很有帮助。你应该达到的基准:

  • Implement a transformer end-to-end端到端实现Transformer
  • Implement causal, cross, and self attention实现因果注意力、交叉注意力和自注意力
  • Implement flash attention实现闪存注意力
  • Implement the attention backward pass实现注意力的反向传播
  • Implement an MLP forward and backward pass实现MLP的正向和反向传播
  • Implement a simple training loop with SGD in PyTorch or JAX使用PyTorch或JAX实现简单的SGD训练循环

If you can do all of these from scratch under time pressure, you're in good shape.如果你能在时间压力下从头完成所有这些,你的状态就很好了。


How I Prepared: Emotional我的准备方法:情感方面

I can't speak for everyone, but for me this process was where my resilience went to die. I've always handled interviews and exams well without any particular strategy... not this time.我不能代表所有人,但对我来说,这个过程让我耗尽了韧性。我一直面试和考试都表现不错,不需要特别的策略……但这次不同。

If you're doing okay emotionally, skip this section. I don't want to plant anxiety where there isn't any.如果你情绪上没问题,跳过这一节。我不想制造不必要的焦虑。

The practical stuff first实用建议先行

My biggest problem was sleep: I couldn't sleep well the night before an interview, which becomes a serious issue when you have 10 interviews in a week. I also couldn't eat, I'd get nauseous at the sight of food before an interview. My solution was to chug litres of Coke for the sugar. I'm not claiming this is optimal, but it's the best solution I could come up with.我最大的问题是睡眠:面试前一晚我睡不好,当你一周有10场面试时,这就成了严重问题。我也吃不下东西,面试前看到食物就觉得恶心。我的解决办法是狂喝可乐补充糖分。我并不是说这是最佳方案,但这是我能想到的最好的办法。

Beyond survival, I would recommend: regular exercise, a consistent evening routine, and not isolating yourself socially. Running before interviews helped me a lot, it burned off nervous energy and reset my head. (If you do this: take it easy and eat enough carbs). During the worst weeks I made a rule to have dinner with friends any evening I didn't have an interview the next morning. It helped me a lot.除了生存策略之外,我建议:规律锻炼、坚持晚间作息、不要孤立自己。面试前跑步对我帮助很大,它消耗了紧张能量,让头脑恢复清醒。(如果这么做:轻松跑,吃足够的碳水化合物)。在最糟糕的几周,我规定如果第二天早上没有面试,就和朋友一起吃晚餐。这对我帮助很大。

The pre-interview ritual面试前仪式

I found a lot of comfort in having a consistent pre-interview ritual. I'd put fresh flowers in my background (I got quite a few compliments on that), do my make-up (it was relaxing to focus on something else for a while — for the guys, maybe try skincare?), and watch the same few comfort videos on YouTube. My rotation: Alysa Liu's ice skating and life teachings, alongside the classic Lord of the Rings, Lord of the Rings, and Lord of the Rings. Email me for more wholesome video recommendations拥有一个固定的面试前仪式让我感到安心。我会在背景里放上鲜花(因此获得不少称赞),化妆(专注于其他事情可以放松——对于男士,也许可以试试护肤?),然后观看几个固定的舒缓视频。我的轮换视频:Alysa Liu的花样滑冰和生活哲理,以及经典的《指环王》系列。如需更多积极视频推荐,请给我发邮件。

The harder part更难的部分

At a certain point my anxiety was holding me back more than my preparation was. My mind would occasionally go blank mid-interview. I genuinely considered starting therapy during the process, but ran out of time before I could. In hindsight, that kind of reflection is more useful before you start (knowing your triggers, your relationship with failure, what your sense of worth is actually tied to) so you're not discovering it under fire like I did. Wouldn't recommend lol在某个时间点,我的焦虑对表现的阻碍超过了准备程度。面试中我的大脑会偶尔一片空白。我真的考虑过在这个过程中开始心理治疗,但还没来得及就结束了。事后看来,这种反思在你开始之前更有用(了解你的触发点、你与失败的关系、你的自我价值究竟依附于什么),这样你就不会像我一样在压力下才发现。不推荐这么做,哈哈。

But this brings me to the thing I most want to say: your worth as a human being is not going to be decided by these interviews (I know I would have rolled my eyes at this a few months ago but it's true!). The process is inherently stochastic, and sometimes the universe has a sense of humour about it. The morning of my DeepMind interview I woke up at 5am in a cold sweat having suddenly remembered I hadn't reviewed topic X. I got my phone and asked Claude to summarise topic X for me, then I went back to sleep. When I joined the interview at 9am, the interviewer said: "What do you know about topic X?". I'm not saying there's a god, but if there is, they were clearly on my side that day. But also, you are allowed to have a bad day. Failing to explain why forward KL is mean-covering and reverse KL is mode-seeking in an interview does not make you a bad ML researcher. I absolutely bawled my eyes out after messing up exactly that question after dealing with forward vs reverse KL in two separate papers. You will mess up, even on things you know, and that's okay.但这引出了我最想说的话:你作为人的价值不会被这些面试决定(我知道几个月前我会对此翻白眼,但这是真的!)。这个过程本质上是随机的,有时候宇宙会跟你开玩笑。我DeepMind面试的那天早上,凌晨5点我突然想起没有复习某个主题X,冒了一身冷汗。我拿起手机让Claude总结主题X,然后继续睡觉。9点面试时,面试官问:“你对主题X了解多少?”我不是说有上帝,但如果有,那天他们显然站在我这边。但同样,你也有权利度过糟糕的一天。在面试中没能解释清楚前向KL是均值覆盖、反向KL是众数覆盖,并不会让你变成糟糕的ML研究者。我在搞砸了那个问题后大哭了一场,而我两篇论文都涉及前向与反向KL。你会搞砸,甚至在你熟悉的事情上,这没关系。

Books that helped有帮助的书籍

Not specific to interview anxiety, but useful for the underlying mindset work: The Now Habit by Neil Fiore, The Gifts of Imperfection by Brené Brown, Mindset by Carol Dweck, and The Tyranny of Merit by Michael Sandel.不特指面试焦虑,但对底层心态工作有用:《现在习惯》(Neil Fiore)、《不完美的礼物》(Brené Brown)、《终身成长》(Carol Dweck)和《精英的傲慢》(Michael Sandel)。


How I Prepared: Logistics我的准备方法:后勤

One interview per day Personally, I preferred it this way, when I could manage it. It's not always possible, but interviews are exhausting and you are naturally going to underperform in your third interview of the day. My rhythm was: do the interview in the morning, then spend the rest of the day preparing for the next one. It helped me avoid feeling like I was constantly context-switching mid-day.每天一场面试。就个人而言,我更喜欢这样,只要时间允许。虽然并非总是可能,但面试非常耗费精力,一天中的第三场面试自然表现更差。我的节奏是:上午面试,然后当天剩余时间准备下一场。这帮助我避免在一天中频繁切换上下文。

Start with companies you care less about. Smaller startups, companies in locations you're not keen on, roles that are interesting but not your top choice. You'll get a feel for what the process looks like, calibrate your confidence, and get a realistic sense of what compensation looks like before you're negotiating for the offers you actually want.从你不那么在乎的公司开始。小型初创公司、你不太感兴趣地点的公司、有趣但不是首选的角色。你会对流程有所了解,校准自信,并在为你真正想要的offer谈判之前对薪酬有现实的认识。

Think about timing. Some companies move fast, others are extremely (and unpredictably) slow. Once a process starts it tends to move at a predictable pace, so if company A sent you a link to a test you can do whenever to start your interview process, and you are waiting to hear back from Company B… wait for Company B to schedule their first interview before doing the test from Company A. The goal is to have offers land in roughly the same window so you have real leverage and real choices. Obviously, the process is, again, very stochastic so timing is hard in practice. For example: I tried to message someone I knew at Anthropic to try and speed things up with them but ultimately I failed to start the process with them before the deadline from Deepmind expired.考虑时间安排。有些公司动作很快,其他公司则极其(且不可预测地)缓慢。一旦流程启动,通常以可预测的速度推进,所以如果A公司发给你一个测试链接,你可以随时开始面试流程,而你在等待B公司的消息……那么等待B公司安排第一次面试之后再去做A公司的测试。目标是让offer大致在同一时间到来,这样你才有真正的筹码和选择。显然,流程再次非常随机,所以时间安排在实践中很难。例如:我试图联系Anthropic认识的人加快流程,但最终还是没能在DeepMind的截止日期之前启动他们的流程。

Tell every company about your other processes. I know it feels uncomfortable for some people but it's completely normal and expected. It keeps timelines clear, encourages processes to chug along nicely, and often prompts companies to move faster if they're interested. I think companies will also consider you as a more serious candidate if they know multiple other companies consider you a serious candidate.向每家公司告知你的其他流程。我知道有些人会觉得不舒服,但这完全正常且是预料之中的。这能保持时间线清晰,鼓励流程顺利推进,并且如果公司对你有兴趣,通常会促使他们加快速度。我认为,如果公司知道其他多家公司也把你视为认真候选人,他们也会更认真地考虑你。


Negotiation谈判

I found this blog post helpful as a starting point, though I'll be honest, I didn't follow its advice. The post recommends treating it like a blind auction and not revealing competing offers. That didn't work for me: several companies explicitly asked for proof of other offers before increasing theirs, and one even questioned my screenshots (lol). So much for the poker face approach.我发现这篇博客文章作为起点很有帮助,但说实话,我没有采纳它的建议。文章建议像盲拍一样对待,不透露其他offer。这对我不起作用:有几家公司明确要求提供其他offer的证明才肯提高报价,甚至有一家质疑我的截图(哈哈)。所以扑克脸方法行不通。

A few things I learned:我学到的一些东西:

Companies can move their numbers significantly if they want you... more than I expected. It's always worth asking. Most companies were open to negotiating.如果公司真的想要你,他们可以大幅调整数字……比我预想的更大。总是值得一问的。大多数公司对谈判持开放态度。

Deadlines varied from one week to two weeks to a vague "take a reasonable amount of time." In my experience companies weren't flexible about extending them (but they were for some friends of mine!), so factor that into your timing.截止日期从一周到两周不等,有的甚至模糊地说“合理时间”。根据我的经验,公司不太愿意延长截止日期(但我的一些朋友遇到了愿意延长的),所以要把这一点纳入时间安排。

Recruiters are surprisingly good at reading you. A few figured out my actual preferences just by asking vague questions, but in fairness I'm extremely transparent and probably very easy to read. Be careful, even small signals matter: how often you mention a company, how you talk about them, all of it gets noted. If a recruiter knows their company is already your preferred choice, negotiating is going to be harder.招聘人员非常擅长读你的心思。有几个人仅仅通过问含糊的问题就看出了我的真实偏好,但公平地说,我非常透明,可能很容易被看穿。要小心,即使是很小的信号也很重要:你提到某个公司的频率、你如何谈论他们,所有这些都会被注意到。如果招聘人员知道他们公司已经是你的首选,谈判会变得更难。

Companies track historical data on candidate choices. If you tell Anthropic you're seriously considering an offer from Peppers Burgers, they have data on how often candidates with both offers actually chose the latter. If the answer is "almost never," your bluff doesn't work. This is also why competing offers from actual peers (OpenAI or another top lab) carry actual weight in a way that other offers don't. But again, it's very difficult to line up these processes.公司会追踪候选人选择的历史数据。如果你告诉Anthropic你正在认真考虑Peppers Burgers的offer,他们有数据知道有多少同时手握两份offer的候选人最终选择了后者。如果答案接近“几乎从不”,你的虚张声势就不管用了。这也是为什么来自真正同行公司(OpenAI或其他顶级实验室)的竞争offer具有实际分量,而其他offer则不然。但同样,很难让这些流程的时间线对齐。


Decision Making Process决策过程

I cannot speak to your situation, but personally I was quite insecure at the beginning of this process and I was tempted to accept some of the early offers I got (rather than letting them expire and keep interviewing) out of fear I wouldn't find anything else. I did find better. Obviously it's impossible to predict how future processes are going to go, but trust your gut.我无法代表你的情况,但就我个人而言,在这个过程开始时我很没有安全感,曾一度想要接受早期的一些offer(而不是让它们过期继续面试),因为担心找不到更好的。后来我确实找到了更好的。显然,无法预测未来的流程会如何发展,但要相信你的直觉。

For choosing between offers: everyone weighs things differently: location, compensation, prestige, type of work, free food and Coke (the last two are very important to me). I had a rough preference ordering before I started, which shifted as I learned more about teams and culture and compensation. My very sophisticated vibe-based ranking system then collapsed entirely when I fell in love with Isomorphic Labs and then DeepMind made me an offer.在多个offer之间选择:每个人权衡的因素不同:地点、薪酬、声望、工作类型、免费食物和可乐(后两点对我非常重要)。我在开始前有一个粗略的偏好顺序,但随着对团队、文化和薪酬的了解加深,这个顺序发生了变化。我那个非常复杂的基于感觉的排名系统在我爱上Isomorphic Labs然后又收到DeepMind的offer时彻底崩溃了。

My solution was to speak to essentially everyone at both companies. In a shocking twist of events, every single person at DeepMind told me they'd choose DeepMind, and every single person at Isomorphic told me they'd choose Isomorphic. Extremely helpful. In the end the most useful thing was talking it through with the people who actually know me (in my case, my boyfriend) and figuring it out from there.我的解决办法是和两家公司几乎所有的人交谈。令人震惊的是,DeepMind的每个人都告诉我他们会选DeepMind,而Isomorphic的每个人都告诉我他们会选Isomorphic。非常有帮助。最后最有用的还是和真正了解我的人(对我来说是我男朋友)讨论,然后从中得出结论。


What I'd Do Differently我会做出哪些改变

Even if the process was successful overall, there is a few things I would change if I was to do this again in the future:尽管整个过程总体成功,但如果将来再做一次,有几件事我会改变:

Keep a spreadsheet. I was convinced I could track everything in my head. Technically yes, but a simple spreadsheet (companies to apply to, where you are in each process, deadlines, contacts) would have stopped me from forgetting to apply to places I was actually interested in. Not rocket science, I know.使用电子表格。我曾相信自己可以在脑子里记住所有事情。理论上可以,但一个简单的电子表格(要申请的公司、每个流程的进度、截止日期、联系人)本来可以防止我忘记申请自己真正感兴趣的地方。这不是什么高深的事情,我知道。

Prepare emotionally, not just technically. The interview process has a way of feeling like a final verdict on your abilities as a researcher and whether your PhD was worth anything. That's not a rational framing, but it's hard to avoid when you're in it. I didn't handle it well, and I think some therapy or at least serious introspection before starting (rather than during) would have helped a lot.不仅从技术上,还要从情感上做好准备。面试过程有时会让你觉得这是对你研究能力和博士价值的最终裁决。这种想法并不理性,但当你身处其中时很难避免。我没有处理好这一点,我认为在开始之前(而不是期间)进行一些心理治疗或认真反思会非常有帮助。

Be more proactive about the companies that ignored me. In hindsight I should have cold-emailed someone at the company, expressed my interest directly, and tried to actually get on someone's radar rather than hoping the application form would do it for me. If you really want to work somewhere and you're not hearing back, do something about it.对忽略我的公司更加积极主动。事后看来,我本应该给那些公司的人发冷邮件,直接表达兴趣,并尝试进入他们的视野,而不是指望申请表能为我做到。如果你真的想去某个地方却没有收到回复,那就做点什么。


Technical Topics技术主题

Here is a list of topics I created before I started interviewing. Personally, I was asked a lot about LLMs and RL, reflecting my background. If you have a diffusion background, expect more questions there. I was asked (in some form or capacity) about pretty much all the topics I studied in at least one interview. So... make sure to cover everything well!

Reinforcement Learning强化学习

  • Q-Learning / TD LearningQ学习 / 时序差分学习
  • Bellman Equations贝尔曼方程
  • PPO
  • GRPO
  • GAE
  • Variance Reduction in RL强化学习中的方差缩减
  • DPO (Direct Preference Optimisation)DPO(直接偏好优化)
  • Policy Gradient Theorem策略梯度定理
  • On-Policy vs Off-Policy同策略 vs 离策略
  • Exploration vs Exploitation Dilemma探索与利用困境
  • Credit Assignment Problem信用分配问题
  • MuZeroMuZero
  • World Models / Dreamer世界模型 / Dreamer
  • AlphaGoAlphaGo
  • Soft Actor-Critic软演员-评论家
  • Model-Based vs Model-Free基于模型 vs 无模型
  • Markov Property马尔可夫性质
  • Monte Carlo vs TD蒙特卡洛 vs 时序差分
  • Actor Critic演员-评论家
  • SARSA
  • Importance Sampling重要性采样
  • Markov Decision Process马尔可夫决策过程
  • Curriculum Learning课程学习

LLMsLLMs

  • Flash AttentionFlash Attention
  • LoRALoRA
  • TransformerXLTransformerXL
  • GriffinGriffin
  • PerceiverPerceiver
  • Scaling Laws缩放定律
  • Mixture of Experts混合专家模型
  • LLM scaling factorLLM缩放因子
  • RoPERoPE
  • Sinusoidal embeddings正弦位置编码
  • Relative positional embeddings相对位置编码
  • LLM vs RNN vs S4LLM vs RNN vs S4
  • Tokenisation分词
  • Pretraining预训练
  • Finetuning微调
  • RLHF
  • Decoding techniques解码技术
  • Causal Attention因果注意力
  • Cross Attention交叉注意力

Generative Modelling生成建模

  • GANs生成对抗网络
  • VAEs and VAE ELBO变分自编码器与VAE ELBO
  • Score Function得分函数
  • Diffusion Forward Process扩散前向过程
  • Diffusion Reverse Process (DDIM / DDPM)扩散反向过程(DDIM / DDPM)
  • Diffusion Forward / Reverse SDE扩散前向/反向SDE
  • Flow Matching ODE流匹配ODE
  • Classifier Free Guidance无分类器引导

Applied ML应用ML

  • Tensor Parallelism张量并行
  • FSDP
  • DDP
  • Pipeline Parallelism流水线并行
  • Communication Primitives通信原语
  • Mixed precision training混合精度训练
  • Gradient checkpointing梯度检查点
  • Gradient accumulation梯度累积
  • Profiling性能分析
  • Gradient clipping梯度裁剪
  • Numerical precision tricks数值精度技巧
  • Exploding / vanishing gradients梯度爆炸/消失
  • Floating point representation浮点数表示
  • JIT compiling即时编译
  • JAX, PyTorch, TensorFlowJAX、PyTorch、TensorFlow

General ML通用ML

  • Curse of dimensionality维度灾难
  • S4
  • CNNs卷积神经网络
  • RNNs / LSTMsRNN / LSTM
  • Autoencoders自编码器
  • Gumbel-SoftmaxGumbel-Softmax
  • MLE vs MAPMLE vs MAP
  • Newton's Method牛顿法
  • Linear Regression线性回归
  • Activation Functions激活函数
  • Loss Functions损失函数
  • No Free Lunch Theorem没有免费午餐定理
  • BatchNorm / LayerNorm / RMSNormBatchNorm / LayerNorm / RMSNorm
  • Variance and Covariance方差与协方差
  • Adam / AdamW / AdagradAdam / AdamW / Adagrad
  • Bias-Variance Tradeoff偏差-方差权衡
  • Backprop反向传播
  • Regularisation Methods正则化方法
  • Unsupervised vs Supervised无监督 vs 有监督
  • Clustering Algorithms (e.g. k-means)聚类算法(如K-means)
  • K-Nearest NeighboursK近邻
  • SVMsSVM
  • BoostingBoosting
  • BaggingBagging
  • Decision Trees决策树
  • Ensembles集成方法
  • Bayes Theorem贝叶斯定理
  • Precision / Recall / F1 / AUC-ROC精确率/召回率/F1/AUC-ROC
  • KL DivergenceKL散度
  • Jensen-Shannon Divergence詹森-香农散度
  • Weight initialisation权重初始化
  • Gradient Descent / SGD梯度下降 / SGD
  • Overfitting / Underfitting过拟合 / 欠拟合
  • Cross validation交叉验证
  • Data Whitening数据白化
  • Convex functions凸函数
  • Early Stopping早停法
  • Domain Adaptation领域自适应
  • Dimensiolity Reduction降维
  • Transfer Learning迁移学习
  • Few shot / Zero shot learning少样本/零样本学习
  • Second Order Methods二阶方法
  • Expectation期望
  • Entropy
  • PDF / PMFPDF / PMF
  • Confidence Intervals置信区间

Linear Algebra线性代数

  • Positive Semi-Definite半正定
  • Jacobian雅可比矩阵
  • Eigenvectors / Eigenvalues特征向量/特征值
  • Hessian海森矩阵
  • Inverse of a matrix矩阵的逆
  • Dot product点积
  • Null space / Image space零空间/像空间
  • Orthogonality正交性
  • Linear independence线性无关
  • Singular matrices奇异矩阵
  • Rank / Span秩/张成空间
  • Determinant行列式