I led product and design for Gusto's AI reporting agent, a natural language interface for querying payroll, money, retirement, HR, and benefits data. For roughly three months, I covered both product and design roles while the PM seat was open.
Initial skill built into Gusto’s help agent, which answered payroll questions
The project started as a skill within Gusto’s help agent, but tool-call reliability exposed limitations in the approach. While iterating on the skill with engineering, I built a vision prototype and validated it directly with users. The insights helped drive leadership alignment around a pivot to a standalone agent within the reporting dashboard, giving us a focused environment to learn, iterate, and scale the experience.
Prototype for a standalone Reporting Agent
Customer testing quickly revealed that users didn’t want AI to replace reporting, they wanted it to remove the friction around it. They preferred a conversational experience that still delivered familiar outputs like tables, PDFs, and CSVs. By letting users describe what they needed, we reduced the effort of finding reports, configuring filters, and manually building queries. We integrated these conversations into existing report history, making AI a natural extension of the reporting workflow.
A full-screen agent experience with structured tables and downloadable artifacts
Customer testing revealed a broader opportunity: users saw the agent as more than a reporting tool. Several used it as a real-time copilot alongside their existing workflows, pulling it up on their phones while running payroll or filing taxes on their laptops.
Expanding the agent’s capabilities through intelligent report routing and data visualization
I partnered closely with engineering to define the quality framework behind the agent, including evals, voice and tone, component rendering, and consistent experiences across conversational and traditional interfaces.
The Gus copilot experience, bringing conversational AI into the report-building workflow
For traditional reporting experiences, testing revealed that customers didn’t want AI to replace the reporting workflow; they wanted it to make the workflow faster. I designed a copilot experience that combined natural language prompts with familiar GUI controls, allowing customers to quickly generate a report and then refine it through filters, sorting, and column adjustments.
This work points toward a future where reporting becomes less about building reports and more about answering questions. By combining company context, account activity, and conversational interactions, AI can proactively surface relevant insights, help customers refine their analysis, and make deeper exploration feel natural within the existing workflow.
I led a pilot exploring how Gusto could automate recurring administrative workflows, laying the foundation for a broader task automation platform where customers delegate routine work to Gusto's AI agents.
The initial task automation for compliance training
We started with compliance training. Instead of asking employers to identify employees, configure enrollments, and track completion, Gusto handled the entire workflow. Customers simply reviewed the proposed action and approved it. The pilot increased task completion, training revenue, and employee engagement, validating that customers were willing to delegate routine administrative work to AI agents.
A spectrum of agent capabilities, from guided interactions to proactive automation
As a next step, I explored different levels of task automation, from conversational workflows that guided users through tasks to fully automated actions with human-in-the-loop approvals. This framework helped shape a reusable system for building AI-powered task automation across product areas.
I secured leadership alignment and roadmap commitment with three app teams. Although the broader rollout was paused during a shift in Gusto’s AI strategy, the compliance training pilot and UXR validated the potential of AI-powered task automation.
I partnered with engineering to design the architecture of Gusto's reporting platform, creating a foundation that could scale with the needs of the business. At the time, reporting relied on disconnected legacy systems that struggled to support new and existing reports, often resulting in slow, unreliable experiences.
The existing reporting architecture
The core idea was a centralized reporting database: a single source of truth powering all reports and designed to support future platform offerings.
A centralized reporting platform designed to scale across products
We partnered with each product team to bring their domain data into a shared pipeline, reducing load on existing systems while improving reporting performance and reliability. On top of that foundation, we built a governed API layer that provided a secure, consistent gateway for new capabilities like widgets and data visualization.
A platform foundation that accelerates future AI experiences
That foundation is what allowed Gusto to move quickly as AI capabilities emerged. With the API already in place, we could build MCP servers on top of it instead of creating new integrations from scratch. The first MCP server powered our initial ChatGPT and Anthropic integrations, launching in under six weeks. When Gusto's Co-Founder experience launched next, reporting was available from day one, and the team shipped an MVP in four weeks. The same foundation now powers our AI reporting agent. The impact compounds over time: as domain teams add new data to the pipeline, the time required to bring new capabilities to customers through AI has decreased from months to days.
Domain-specific data views enabling more accurate agent responses
As AI capabilities evolved, we uncovered a new architectural challenge: giving the LLM raw access to Gusto’s database wasn’t enough to reliably join data and return accurate answers. We introduced standardized "views" that created stable, domain-specific layers for data like payroll summaries, expenses, and team member information. This structure improved data integrity, reduced errors, and helped the agent deliver faster, more reliable answers.