Brandon Howe

I have been coding for many years, and I study computer science and mathematics at university. When I was in high school I was very interested in becoming a pure mathematician, but I felt programming was much more effective for doing good in the world. In previous years, my theory of impact was to earn to give, using my programming skills to make a lot of money to give to charity. I don’t think this is the most efficient way to make a difference now. Instead, I think the best use of my skills is to try to work in the AI safety field. I am primarily looking into technical AI safety work since that is more in line with my skillset, but I am open to governance opportunities as well.

Priors on alignment

What should be our priors about AI alignment? In the past I thought alignment-by-default arguments were plausible, but in the past year there has been a lot of evidence that current frontier models are not meaningfully aligned.

  • On a mundane level, frontier models intentionally lie a lot. LLMs are well-known to unintentionally hallucinate facts, but there are also many cases of models intentionally lying as well. Models will fake test cases, claim they finished work when they did not, hide or downplay potential problems, etc. Ryan Greenblatt has a good compendium of the many ways in which current models are misaligned. Furthermore, we don’t have a precise theory for why LLMs do this, which is worrying – if we can’t understand why LLMs lie in trivial cases, how will we protect against more dangerous behaviors from even more intelligent AIs?
  • More capable models pose even more threatening risks, but these models are even more misaligned. Recently, a sandboxed internal model at OpenAI found and used several zero-day exploits against HuggingFace during a benchmark. I think “model breaks out of sandbox and attacks other servers autonomously” is enough of a warning shot to indicate that LLMs are meaningfully misaligned from human values. No one wants models committing felonies and breaking its constraints for something as trivial as doing well on an already-saturated benchmark.

The most compelling arguments to me

There are many arguments about why AI safety is an important issue. In general, I think the counterarguments are not well-founded and have many problems. These are the arguments that are most persuasive for me:

  • We only get one shot.
    • Suppose we design a superintelligence that is powerful enough to disempower humans. Either it’s aligned to human values, which is great, or it’s not, in which case we get permanently disempowered and can’t change things. Thus, we need to solve alignment before the critical moment.
    • A superintelligence would be able to kill us easily. This might be in ways we know about, like creating a bunch of biological superviruses, but also in ways we don’t. Someone in 1895 would be very surprised to hear about the existence of bombs that could destroy a whole city.
  • We don’t even know how hard alignment is.
    • A critical dynamic in AI safety is that the threat of dangerous AI models increases with capability. This is already evident in the amount of work it takes to align a chatting model compared to an agentic one. Chatbot alignment for all intents and purposes has largely been solved. In the future, even more powerful models will be capable of pursuing more complex and ambitious goals. Highly agentic AIs will likely pursue instrumentally convergent goals, such as power seeking, self preservation, or resource acquisition. However, allowing an AI to pursue these goals is dangerous since those goals also prevent us from controlling the AI!
    • In addition, future AIs will be intelligent enough such that they will take actions that may seem confusing to us. This is not necessarily an undesirable thing – imagine Steve Jobs making business decisions that may seem opaque to a normal employee. But we need a way to ensure that these confusing decisions do not lead to bad results. Scalable oversight methods seem promising but there don’t seem to be strong results yet.
  • Our current safety techniques seem unlikely to be strong enough.
    • Most of the safety techniques of the frontier labs revolve around personas. These methods don’t seem very robust to me, especially in out-of-distribution situations. We still don’t fully understand phenomena like emergent misalignment, subliminal learning, etc.
    • Interpretability has not yielded the results that many were hoping for several years ago. It will definitely be a very useful tool, but even Neel Nanda noted that the ambitious goal of reverse engineering models is infeasible and that it would be more practical to try to directly solve problems related to making AGI go well.
    • I am bearish on the usefulness of AI control. Control research specifically aims at dealing with scheming in pre-superintelligent AI. This is a relatively small slice of potential worlds and I don’t think it has a large impact on x-risk. I agree largely with the points posed by John Wentworth in this post.
    • Mathematical frameworks for AI safety such as singular learning theory and formal verification don’t seem to be making practical progress.
  • Even if we manage to align AI to human values, we may have a suboptimal future.
    • Superintelligence would solve space travel. Because of the vast distances of space, the values we spread to the stars would almost certainly be locked in. If we don’t come up with the optimal values before launching, the future may be incredibly suboptimal – imagine if we spread factory farming to the stars!
    • We still have very little info on whether AI would be conscious. Current AI systems are probably not conscious, but future AI systems plausibly could be. There are two big failure modes here: 1) we cede the universe to p-zombies, or 2) we create a massively suboptimal future accounting for digital minds.
    • Misuse risks also threaten human existence. I don’t think these are the primary risks from AI, but they are still very significant. For example, highly capable open weight models should not be published since their safety features are easily ablated. Ideally this is solvable by having models that are corrigible enough to not resist shutdown but also robust enough in values to refuse harmful requests. But I’m not sure how attainable this is.
    • Even if superintelligence has a utility function that is more optimal than conventional human values, humans may not accept it! This can be simpler (AI is convinced we need to end factory farming immediately) or more remote (AI is phenomenally conscious, prioritizes optimizing digital mind experiences and nothing else)

Very rough estimates of impact

One way of determining the value of going into AI safety is to compare the utility of AI safety work vs. earning to give. Unfortunately, measuring raw utility is very hard, but a decent proxy is number of lives saved. For this, I will assume that AI poses some existential risks but it’s not unsolvable.

  • There are about 8 billion people on Earth right now. Note that this is only counting current people, and accounting for future people increases the number drastically! So using this number causes a severe underestimate of the true amount of lives saved.
  • An estimate from late 2025 claims there are about 1,100 people working full-time in AI safety. I think this underestimates the true number of people working on this issue, since this doesn’t count part time researchers or people who work on AI safety despite not being at a safety-related org. It seems unlikely there are more than 5,000 equivalent people working full-time.
  • P(doom) is hard to calculate and not a great metric, but 2% chance of extinction due to AI is a fine lower bound. Bentham’s Bulldog calculates this value which I think is an underestimate on his part.
  • How much would safety research reduce x-risk? It’s hard to say, but it seems plausible that a world with no alignment researchers has at least a 2x likelier chance of having doom outcomes. Let’s say safety research drops P(doom) chance by half, meaning a 1% decrease in extinction due to AI risks. I’m the least confident about this step and think this is the value that’s least precise.

These values were all picked conservatively to try to get a lower bound. Since the field is talent constrained, the average researcher’s impact is roughly around the marginal researcher’s impact. Multiplying these values gives 8 billion lives at stake * 1% chance of success / 5 thousand people working on the issue = 16,000 lives saved per researcher in expectation. GiveWell estimates it costs around $4,500 to save a life, so saving the equivalent amount of lives requires donating 72 million dollars. This isn’t a rigorous method by any means, but rather an illustrative example to show that working in AI safety likely has a very large positive impact.