shipgremlins. Star on GitHub

OPEN SOURCE. SLIGHTLY FERAL.

AI product managers
that actually
use your app.

Give each PM a mandate. They explore in a real browser, find the gaps, and help turn approved tickets into tested fixes. You keep building. They keep improving.

Self-hosted · Real browsers · Your final say

MEET THE NIGHT SHIFTGREMLIN / 001
A mischievous lime green gremlin holding a browser bug report
“Found something.”Of course you did, little buddy.
001
Will absolutely click that button twice.
EARLY ALPHA / BRING YOUR OWN CURIOSITY

Adopt your first gremlin.

Node 22.12+ · Setup guide
FROM SOURCE
$ git clone https://github.com/AgentBurgundy/shipgremlins.git
$ cd shipgremlins && npm ci
$ npm run setup -- --check
BUILT AROUND YOUR WORKFLOW
◈ GitHub
▲ Vercel
◒ Linear
⌘ Playwright
More on the roadmap

A LITTLE CHAOS.
A LOT OF SCREENSHOTS.

Your app's
new night shift.

Give them a mandate.
Let them poke around.
Review what they find.

01 / MEET THE CREW

A particular set
of nitpicks.

One PM owns one area. Give each a mandate, a schedule, and boundaries. Let curiosity do the rest.

THE SUSPICIOUS ONE
Security gremlin inspecting a shield and lock with a magnifying glass

The Security PM

“Should that user really be able to do that?”

  • Roles and permissions
  • Cross-tenant boundaries
  • Auth and session edge cases
THE NOSY ONE
Feature gremlin examining an unusually long CSV spreadsheet

The Feature PM

“Okay, but what if I upload this?”

  • Deep feature exploration
  • Awkward inputs and empty states
  • Acceptance and regression checks
THE PICKY ONE
Experience gremlin inspecting a button on a mobile phone

The Experience PM

“It works. But does it make sense?”

  • Signup and first-run friction
  • Mobile layouts and accessibility
  • Confusing flows and dead ends

Not fixed job titles. Start with a template, then write the mandate your product needs.

02 / SHOW YOUR WORK

“Trust me”
isn’t a test.

Every useful finding needs a reproduction. Every verified fix needs evidence. Follow the trail from a browser observation to a production merge.

Illustrative workflow. No live agents are running on this page.

MANDATEKeep every tenant in its lane.EXAMPLE RUN
01
OBSERVE

A member opens another team's export.

Reproduce with two isolated test accounts. Record the request and the actual browser state.

02
PROPOSE & APPROVE

One finding. A clear acceptance test.

Evidence goes into Linear. Approved scope becomes a developer task.

03
IMPLEMENT & VERIFY

The owner can export. The other team can't.

Check the deployed revision, capture screenshots, and rerun the regression.

Verified is a milestone. Production is Done.Staging review stays in your hands.

03 / GUARDRAILS INCLUDED

Ambitious agents.
Sensible boundaries.

Your crew can be relentless without having the keys to everything.

01

A mandate, not a blank check.

Control the scope, approved tickets, test environment, and amount of work in flight.

02

Break. Diagnose. Repair.

Bounded retries and a dedicated recovery lane keep ordinary work paused while a broken integration is repaired.

03

No evidence, no promotion.

Require current deployment evidence before a promotion opens. Stale checks and a confident comment don't count.

04 / BRING YOUR STACK

Your tools.
A few new teammates.

Keep your code, tickets, and infrastructure.
Add a crew around them.

AVAILABLE IN THE ALPHA

GitHub + Vercel

GitHub Actions runs the agents. Vercel hosts the previews. Linear tracks the work.

Self-hosted runners · Playwright MCP · Claude Code
NEXT ON THE ROADMAP

GitLab + Railway

The same loop, with GitLab CI and an isolated Railway test environment.

Provider adapters · More AI runtimes · GCP certification

THEY'RE SMALL. THEIR STANDARDS AREN'T.

Put a little
gremlin to work.

Start with the source. Inspect your setup.
Give your first PM a job.

Get ShipGremlins on GitHub
YOUR TERMINALNode 22.12+

# Bring the crew home

git clone https://github.com/AgentBurgundy/shipgremlins.git
cd shipgremlins
npm ci

# Check your tools and connections

npm run setup -- --check
Read the setup guide

Early alpha, built in the open. Bring your own provider credentials. A full management dashboard and additional provider stacks are on the roadmap.

A FEW FAIR QUESTIONS

Before you let
them loose.

Does it ship to production by itself?

No. Agents work on approved scope in a test environment. Promotion has verification gates, and the owner retains staging and production merge decisions. Reconciliation marks work Done only after production inclusion is proven.

Can I run it on my own server?

Yes. The alpha runs through GitHub Actions with your self-hosted runners. The setup command checks prerequisites and initializes configuration. Containers are included for the CLI and local site; they are not an agent control plane.

Which AI models can I use today?

The current workflows use Claude Code. Additional runtimes are planned. You control the credentials and model configuration in your own environment.

Does it work with any web app?

The browser can explore many kinds of apps, but reliable testing needs configured accounts, isolated data, and acceptance criteria. The first integration is GitHub with Vercel; GitLab and Railway are next.

Can it create test images and CSVs?

Yes. The fixture CLI generates CSVs from JSON and reproducible PNG test images, ready to upload through Playwright MCP. Semantic AI image generation and automatic test-data cleanup are on the roadmap.