Inspiration
The inspiration for this work came from the Assistant Axis paper by Anthropic from early 2026. Additionally, investigating underrepresented issues in AI safety with an interpretability element was a huge inspiration also. This work represents an approach to uncovering internal language representations differences through an interpretability lens.
What it does
The project helps build an investigative framework to understand how different languages are represented within a single model. Ultimatily, this work should reveal what types of improvements are realistic and appropriate to ensure beneficial and safe AI for everyone.
How we built it
We heavily relied on existing source code of the original Assistant Axis paper and built additional code to help run experiments, analyze and visualize results.
Challenges we ran into
We ran into several practical challenges w.r.t. compute, design decisions, formatting data and data analysis and interpretation.
Accomplishments that we're proud of
We are proud of having accomplished our work and offering a hopefully new perspective on conducting research within this in our opinion underrepresented research field.
What we learned
We learned that LLM multilinguality is a complex subject and the methods in which this concept is approached may not be the most fair nor safe.
What's next for Understanding Multilinguality in LLMs
There's a lot to uncover here. We mention a number of additional research directions that we would like to further investigate, including different models, more languages, understanding model training differences and using a different interpretability space (traits instead of roles). Due to the nature of this research, particularly the use of AI-as-a-Judge, every methodological choice opens new avenues for future work.
Built With
- interpretability
- python
- substack


Log in or sign up for Devpost to join the conversation.