Conversational Browser Agent
Natural language in. Visible browser actions out.
A conversational agent that operates a real Chrome session, asks for missing context, and streams step-by-step browser screenshots to a React interface.
The problem
Browser workflows are often fragmented and difficult to automate reliably when authentication, changing selectors, and missing user information interrupt the happy path.
How it comes together
An asynchronous conversation loop coordinates intent, browser state, and Playwright actions. Screenshot events travel over WebSockets so the person using the agent can follow the browser's progress.
Follow the flow.
Multi-turn collection of missing task information.
Live screenshots after browser interactions.
Two-factor approval waiting and continuation.
Password masking in chat and fallback selector strategies.
Why this approach?
Keep the browser interaction visible. The documented Gmail flow uses the web interface, making its actions inspectable through the same surfaces a person would use.