All projects
PROJECT NOTES / AGENTIC AI

Conversational Browser Agent

Natural language in. Visible browser actions out.

A conversational agent that operates a real Chrome session, asks for missing context, and streams step-by-step browser screenshots to a React interface.

PythonPlaywrightReactWebSocketsOpenAI
AGENTIC AI↗
01Conversation
02Task context
03Playwright
04Chrome
SYSTEM ARCHITECTURE
01 / THE CHALLENGE

The problem

Browser workflows are often fragmented and difficult to automate reliably when authentication, changing selectors, and missing user information interrupt the happy path.

02 / THE APPROACH

How it comes together

An asynchronous conversation loop coordinates intent, browser state, and Playwright actions. Screenshot events travel over WebSockets so the person using the agent can follow the browser's progress.

03 / UNDER THE HOOD

Follow the flow.

SIMPLIFIED ARCHITECTURE
01Conversation
02Task context
03Playwright
04Chrome
05Screenshot stream
04 / BUILD HIGHLIGHTS

Multi-turn collection of missing task information.

Live screenshots after browser interactions.

Two-factor approval waiting and continuation.

Password masking in chat and fallback selector strategies.

THE ENGINEERING CHOICE

Why this approach?

Keep the browser interaction visible. The documented Gmail flow uses the web interface, making its actions inspectable through the same surfaces a person would use.

Explore the project’s supporting sources.Project repository
KEEP EXPLORING

SignalOps · Critical IMS

Vijayshree Vaibhav
YOUR WAY INTO MY WORLD

Vijayshree’s guide.

Projects. Experience. The story behind them.

Portfolio library + Gemini when available
PORTFOLIO GUIDE

Hi! I can help you explore Vijayshree’s projects, experience, and certifications. What would you like to know?

Portfolio sources only · Gemini may process your question.