Inspiration

The inspiration for this work came from the Assistant Axis paper by Anthropic from early 2026. Additionally, investigating underrepresented issues in AI safety with an interpretability element was a huge inspiration also. This work represents an approach to uncovering internal language representations differences through an interpretability lens.

What it does

The project helps build an investigative framework to understand how different languages are represented within a single model. Ultimatily, this work should reveal what types of improvements are realistic and appropriate to ensure beneficial and safe AI for everyone.

How we built it

We heavily relied on existing source code of the original Assistant Axis paper and built additional code to help run experiments, analyze and visualize results.

Challenges we ran into

We ran into several practical challenges w.r.t. compute, design decisions, formatting data and data analysis and interpretation.

Accomplishments that we're proud of

We are proud of having accomplished our work and offering a hopefully new perspective on conducting research within this in our opinion underrepresented research field.

What we learned

We learned that LLM multilinguality is a complex subject and the methods in which this concept is approached may not be the most fair nor safe.

What's next for Understanding Multilinguality in LLMs

There's a lot to uncover here. We mention a number of additional research directions that we would like to further investigate, including different models, more languages, understanding model training differences and using a different interpretability space (traits instead of roles). Due to the nature of this research, particularly the use of AI-as-a-Judge, every methodological choice opens new avenues for future work.

Built With

  • interpretability
  • python
  • substack
Share this project:

Updates

Submission history