Emergence of Instrumental Convergence (IC) properties under Model Post-training
Published:
(This project was done as part of the Bluedot Technical AI Safety Sprint. They’re cool, check them out!)
Published:
(This project was done as part of the Bluedot Technical AI Safety Sprint. They’re cool, check them out!)
Published:
Large language models (LLMs) are often evaluated on clean text. However, real-world inputs are rarely clean. User prompts may contain typos, OCR artifacts, formatting errors, incorrectly copy/pasted fragments, or shuffled and partially corrupted content. These perturbations can affect model outputs in ways that are difficult to predict from input alone.