Is DPO bidirectionally shallow edit?

As part of the BlueDot AI Safety project programme, I studied the bidirectionality of Direct Preference Optimization (DPO), a preference-optimization method used to align models toward desired behaviour. The motivation came from discussions in the Technical Safety course, which drew me toward alignment and interpretability. A paper by Lee et al. (arXiv:2401.01967) studies whether DPO’s edits are “shallow” when used to remove toxicity, and shows that DPO doesn’t erase the toxic capability — it learns a distributed offset that bypasses it, leaving the behaviour reactivatable. ...

Mech interp puzzle

I recently solved the bluedot mech interp puzzle here https://bluedot.org/puzzles/technical-ai-safety?utm_souce=bluedotcommunity My solution was accepted as a correct submission so I am going to describe my approach here.

Do AI models cheat?

Reward Hacking Reward Hacking is exactly what is sounds like, learning a surrogate way to gain reward which doesn’t co-relate with the expectation. The title of the article is do AI models cheat and which shall see what that means and some scenarios where this could be dangerous. AI models are getting larger and smarter and by some estimates smart enough to replace humans at certain jobs? The debate of whether we will get there in the next 5 years or the 50 years is not worth to me but what is worth is as AI models gains human intelligence at certain task can they also gain sub-conciousness? Can they perform some hidden agenda, like saying something but doing something else underneath. ...

Why is AI Safety Challenging?

Safety for weak capability? We are postulating that in near future AI capabilities will grow so much, that it will be superior to human intelligence. In such a reality, enforcing some mechanism that allows AI to still work for the progress of humanity will be necessary. I have been reading the book “Superintelligence” and an opinion that exists among researchers and thinkers, is that we have once chance to build that level of AI that surpasses human intelligence and when we get there, if we don’t do it in a safe way, that could possibly mean in worst case end of humanity. ...