Today, many modern AI systems operate as “black boxes,” meaning researchers cannot fully explain how these systems produce specific outputs.
Interpretability research seeks to understand these internal processes, but the field is still in its early stages, and even experts often cannot clearly explain why a model gives a particular answer. Despite this limited understanding, advanced models like GPT-5 and Claude Opus have already been released to the public.