Hard limit of agent-browser skill recording browser video: every step needs screenshot feedback before next action
Slow interaction, no mouse trail, no typing process
Want to record a feature demo video—driving browser automation with LLM vision is a poor fit, or at least current solutions all suck
