Inspiration
AI Safety is the burning topic right now so picked up this. We're using agents like crazy coming from different sources some are vetted some arent'. When you don't have access to source codes, how do we know it has some hidden agenda or incentives that can impact their behaviour over time? We're addressing it by building a game where we put agents under imaginary court room set up and have agents interact and come up with their own decisions based on given set of rules, hidden agenda, incentives and actions. We also test Game theoretic approach of Prisoner's dilemma where we see Agents at times do selfish confession when they're put under seperate room for investigation.
How we built it
We built it using Claude code in Herdr
Challenges we ran into
Concept was bit challenging to create and convey within teammates and we almost ran out of time building it.
Accomplishments that we're proud of
Proud that we completed the project and submitted on-time. Importantly, we learnt a lot and proved that Agents do behave differently when they're optimized for a variable that you do not control.
What we learned
Agents can lie or deceive when they're sentient. If Robots are going to be desgined by Humans, then human-biases will be part of these Robots that impact everyone's life (Except you're Elon Musk!)
What's next for AMASEI
- Run it 1000 times and get an average AI safety metrics across different experiments and parameter settings
- Can be extended to become a AI safety & Ethics Gaurdrailes as Industry standard
- Have other agents run through our set up and have pass the gaurdrailes
- Help industry develop AI safety use cases as Games
Built With
- claude
- code
- herdr
Log in or sign up for Devpost to join the conversation.