We constantly came up against the same issue again and again — whenever we attempted to automate some web activity, such as filling in a form or extracting data from a dashboard, the script would fail within a few weeks since the website altered its layout. The tools we were using, such as Selenium and Playwright, required us to rewrite all the CSS selectors and XPaths once every small change to the user interface. We then looked into more recent AI browser agents, believing they might overcome this problem, but most of them work by taking a screenshot of the page and then asking a vision model to decide where to click, a method which is slow, expensive, and not very reliable — if each individual step only succeeds 85% of the time, a ten-step task will only succeed once in five attempts. That is why we decided to create Pravah AI, an open-source browser agent that allows any AI model to navigate a website as a person would by simply following plain and simple instructions, without the need for selectors or constant modifications.
Rather than taking a screenshot, Pravah reads the Accessibility Tree of Chrome directly, just as screen readers do, so that the AI receives a clear list of all the buttons, inputs, and links rather than having to work out the information from an image. We also connect directly to Chrome via the Chrome DevTools Protocol rather than going through an extra layer such as Playwright, which enables us to obtain exact click positions, stable IDs that remain unchanged even when the page is altered, and correct handling of nested iframes, something that most other tools have difficulty with. One of our favourite features is action caching – after Pravah has carried out a task using AI, it remembers the precise steps and can replay them later without needing any further AI calls, only making a new AI call if the page actually changes. It proved to be more difficult than we anticipated to get the Accessibility Tree to function reliably on different websites, since a great many sites fail to properly label their elements, and dealing with cross-origin iframes required a lot of additional work because of browser security restrictions. Throughout all this, we have come to the conclusion that structured, text-based information is better for AI than screenshots, and that the real challenge with agents is not getting them to work once, but getting them to work in the exact same way every single time.
Log in or sign up for Devpost to join the conversation.