Eyes Interactive Tutorial Guide
Short on time? Everything the interactive tutorial covers lives on this page: every component, every action, and every takeaway. Read it in about ten minutes, come back to it as a reference, or use the last section to run the whole flow on your own app.
The interactive tutorial walks you through your first visual test with Applitools Eyes: you take a Playwright test that checks a whole page in one line, run it on every browser and device at once, and watch it catch a bug that a list of assertions would have missed. This guide covers the same ground in writing.
Contents
- Meet your workspace
- The flow at a glance
- Chapter 1: Write tests faster
- Chapter 2: One test, multiple environments
- Chapter 3: Visual AI finds the real bug, not the noise
- Chapter 4: Auto maintenance
- Chapter 5: AI closes the loop
- Take it to your own app
- Works with what you already use
- Best practices recap
- Learn more
Meet your workspaceโ
The tutorial simulates a working IDE, so everything you see maps to something you already have. There are five pieces, and only one of them is new.
| Piece | What it is | In the tutorial |
|---|---|---|
| You | You click, pick, and approve. Nothing runs without you. | Buttons and prompts in the chat |
| Your AI assistant | The coding assistant you already use: Cursor, GitHub Copilot, Claude Code, VS Code, and others | Simulated, so it is free and fast. No real tokens are used |
| Eyes MCP | The bridge. It gives your assistant Applitools tools: add visual checks, configure browsers, read results. Everything it does can also be done by hand; the MCP removes the manual work | Installed with one command in Chapter 1 |
| Your IDE | Where the code lives | Embedded as two tabs: Website (the demo Wikipedia app and its HTML) and Code Editor (the Playwright test) |
| Eyes dashboard | A separate Applitools web app holding screenshots, baselines, and batch results. It is not part of your IDE: in real life it lives in your browser | Shown as the Test Results tab: an embedded view of the real Eyes dashboard, pre-loaded for you |
One rule of thumb carries the whole tutorial: you talk to the assistant, the assistant uses Eyes MCP, and Eyes MCP does the testing work you would otherwise do by hand. The dashboard is where you look when you want the evidence.
A few things to notice as you go:
- The progress bar ("Your journey") at the top of the chat tracks the five chapters.
- Tool chips in the chat show every Eyes MCP call by name, so you learn what the server actually does while you watch it work.
- Locked tabs unlock as the story needs them: the Code Editor opens in Chapter 1, Test Results after your first run.
- Highlighted elements: when the tutorial pulses part of the screen, clicking it is how you advance. You are clicking the real dashboard, not a picture of it.
The flow at a glanceโ
The whole tutorial is three runs of the same test, and the story is in the difference between them:
- Run #1 photographs three pages across four environments and saves those 12 pictures as the baseline: the reference for what "correct" looks like.
- Run #2 happens after the app changes: one change on purpose, one by accident, and one that only looks like a change. Eyes shows you which is which. You accept the deliberate one and leave the accident flagged.
- Run #3 proves the accident is fixed and nothing else moved. 12 of 12 pass.
Chapter 1: Write tests fasterโ
The idea. A locator assertion covers exactly the one thing it names. Covering a page properly that way means writing, and then maintaining, dozens of assertions per page, and they still miss everything nobody thought to assert on. One visual check covers the whole page instead: layout, position, size, spacing, color, images, fonts, and how text actually renders, including every element nobody thought to name.
What you doโ
- Install Eyes MCP. One command, shown with a copy button:
npx -y @applitools/mcp@latest
The chat confirms the connection, and from here on the โ tool chips are how you know Eyes MCP is working.
-
Open the Code Editor tab and look at the existing Playwright test. The tutorial highlights every
await expect(...)line: the assertion list you currently maintain. -
Click "Replace the assertions with an Eyes check." The assistant swaps the assertion block for a single call:
import { test } from '@applitools/eyes-playwright/fixture';
test('Wikipedia portal', async ({ page, eyes }) => {
await page.goto(url);
await eyes.check('Wikipedia portal', { fully: true });
});
- Click "Add Eyes checks to the other pages." Three pages, three files, and each file is the same shape: go to the page, check the page.
What that one line doesโ
fully: truecaptures the entire scrollable page, so the footer is checked as carefully as the header. No scrolling code.- The check name (
'Wikipedia portal') is just the label results are filed under. - Note what is not in the test: no locators, no selectors, no class names, nothing that breaks when a developer renames something.
โ Takeaway In the demo suite, 68 functional assertions just became 3 visual checks: about 96% less code and locators to write and maintain, for as long as this test exists. That's hours of work saved after every UI change, functional and visual coverage in one check, and less time reviewing test code an LLM generated for you. And every one of those assertions was something a person had to think of in advance; the check catches the things nobody thought of.
Good to know. Nothing here is Playwright-specific. Eyes has SDKs for Playwright, Selenium, Cypress, WebdriverIO, and native mobile, and the same one-line check works in whatever framework your team already uses. There is nothing to rewrite.
๐ Learn more: Applitools MCP server ยท Playwright quickstart
Chapter 2: One test, multiple environmentsโ
The idea. Your users don't all use Chrome on a laptop. Normally, covering Firefox, Safari, and mobile means running the suite again for each one, on a device farm you have to sign up for and maintain. The Ultrafast Grid works differently: it captures your test once, then renders that capture in parallel on every browser, device, OS, and viewport you pick, in the Applitools cloud. No re-running, no grid to host.
How the Ultrafast Grid works. When your test runs on your machine, in one browser, Eyes doesn't just take a screenshot. It captures the page itself: the DOM, the CSS, and the resources it loads. The Ultrafast Grid then takes that capture and renders it on real browser engines in the Applitools cloud, in parallel, on every environment you configured: desktop browsers (Chrome, Firefox, Safari, Edge) at the viewport size you set, and mobile devices at their native screen sizes in portrait or landscape. Because the navigation runs once, and not on each device, the process is significantly faster and more stable than re-running the test per environment, and there is no device farm, Selenium grid, or QA lab to maintain. Adding an environment is one line of config, not another test run.
What you doโ
- Pick your environments. The chat shows a browser picker. Chrome and iPhone 14 Pro are always included; Firefox and Safari are one click each. Your selection becomes a
browsersInfoblock inplaywright.config.ts:
eyesConfig: {
appName: 'Wikipedia demo',
type: 'ufg',
browsersInfo: [
{ name: 'chrome', width: 1280, height: 800 },
{ name: 'firefox', width: 1280, height: 800 },
{ name: 'safari', width: 1280, height: 800 },
{ iosDeviceInfo: { deviceName: 'iPhone 14 Pro', screenOrientation: 'portrait' } },
],
}
That block is the only place the environment list lives. One config, four environments, still one test.
-
Click "Run the test." Because this is the first run, Eyes has nothing to compare against yet, so it saves what it sees as the baseline: the reference picture of "correct", per checkpoint, per environment. Everything comes back Passed, and that is expected.
-
Open the batch in the dashboard. The result card links straight to it. Confirm three things: 4 environment rows, 12 checkpoints, all Passed. Click into the iPhone 14 Pro row: that is a real mobile rendering of a capture taken once.
Vocabulary the dashboard usesโ
- A batch is one run of your suite.
- Each row is a test in a specific environment.
- Each cell is a checkpoint: the screenshot Eyes captured for one check.
- A baseline is the reference image a checkpoint is compared against. Baselines live in your Applitools account, not in your repo.
โ Takeaway One test. 4 environments. 12 baselines. Eyes rendered your single run everywhere at once: no re-running, no device cloud sign-up, no grid to touch. From now on, every run checks every environment.
- About 10ร faster than running per browser, so CI stops waiting on cross-browser coverage
- No secure tunnels to punch through your firewall
- No timeout or concurrency handling to write: name the devices and browsers, the SDK does the rest
Best practices.
- With the Ultrafast Grid, keep a single Playwright browser project. If you keep multiple local projects, they multiply by grid environments and every result appears twice.
- Treat the first run as a deliberate act: open the batch and look at what was captured before you accept it as the definition of correct.
๐ Learn more: Ultrafast Grid ยท Configuration options ยท Batching
Chapter 3: Visual AI finds the real bug, not the noiseโ
The idea. Pixel diffing finds everything that changed. Visual AI finds what matters. This chapter is the heart of the tutorial.
What you doโ
-
A design request lands. A ticket card appears in the chat: "Increase the Wikipedia logo by 1.2ร on all pages." In a real project this could come from Jira, Slack, or a teammate; the flow is the same. Click "Make this change" and the assistant edits the app's CSS. A one-line change any of us would ship on a Friday, and nobody is going to re-test mobile for a logo resize.
-
Click "Rerun the test on the same environments." Same 3 checks, same 4 environments, now compared against the baselines from Run #1. The result: 9 of 12 checkpoints unresolved. Unresolved is Eyes' honest third status: a visual difference that needs a human decision, bug or intentional change.
-
Click "Inspect the batch result." Instead of clicking through 12 screenshots, the assistant reads the batch through Eyes MCP and reports back in plain language:
- Real change: the logo is about 1.2ร larger on every page that shares the stylesheet. 8 steps, same underlying change everywhere. Not 8 separate issues.
- Possible regression: on iPhone 14 Pro only, the Wikipedia portal's search button is clipped. It is not part of the design request.
-
Open the (view) link and switch to the pixel view (Exact). The whole page lights up: the logo, plus every line of navigation and body text that shifted by half a pixel when the logo grew. This is what any screenshot-diff tool would report. Dozens of failures for one intended change. This is the noise that makes teams stop trusting visual tests.
-
Click "Back to Visual AI (Strict) view." Same page, same pixels. One highlight instead of dozens: the logo. The half-pixel shift is gone from the report because it isn't a defect, and the clipped mobile button is still flagged because it is.
Match levels: how picky should a check be?โ
A match level is the rule Eyes uses to decide what counts as a difference. It is a per-check decision, not a global setting:
| Level | Flags | Reach for it when |
|---|---|---|
| Dynamic (default for new tests) | Like Strict, but auto-recognizes dynamic content (dates, IDs, counters, currency) and ignores the displacement it causes | Almost always. Real pages have content that changes on every load |
| Strict | Anything a person would notice: text, color, layout, size | Brand, legal, forms; anywhere exact rendering is the point |
| Layout | Structure only. Content changes are ignored, but a vanished, overlapping, or resized element is still flagged | Search results, feeds, anything data-driven |
| Ignore Colors | Content and layout; color-only drift is tolerated | Mid-rebrand, themed surfaces |
Set it per check in code:
await eyes.check('Search results', { fully: true, matchLevel: 'Layout' });
โ Takeaway Pixel diff, like other tools, found everything that changed. Applitools Visual AI found what matters: the change you asked for, and a real bug you didn't. Fewer false alarms means every flag is worth opening, and match levels like Strict, Layout and Dynamic let you tune that line per check. And it is deterministic: same input, same result, every run. No guessing.
Best practice. Start on Dynamic for a real application. Tighten to Strict where exact rendering matters. Loosen a region to Layout when its content is data-driven but its structure isn't. Reach for an Ignore region last: it is the only option that makes you blind to that area entirely.
๐ Learn more: Match levels ยท Hide displacement diffs
Chapter 4: Auto maintenanceโ
The idea. Low-maintenance testing that scales with the team. The logo change is intended, so the new screenshots should become the new baseline. With most tools you would now accept it on every page and every environment, one by one: 8 times here, hundreds of times on a real suite.
What you doโ
-
Accept it once. The tutorial highlights the Accept button on one logo checkpoint, in the real dashboard. Click it.
-
Watch what Eyes does with the others. The other logo checkpoints flip to Accepted one after another, each tagged Auto-maintained. Eyes recognized the same difference on the other pages and environments and applied your decision there too.
-
Notice the one that stays unresolved. The Wikipedia portal checkpoint on iPhone doesn't flip, because its diff contains something you haven't ruled on: the clipped search button. That is the difference between Auto-Maintenance and a bulk-accept button. Your decision propagates only where the difference matches, and is withheld where it doesn't.
- Save the batch. The tutorial highlights the Save button. Saving is what turns decisions into baselines: accepted checkpoints become the reference for the next run, and anything left flagged keeps failing until the code changes. Until you save, nothing has moved.
โ Takeaway One click, 7 more resolved on their own. That's Auto-Maintenance: no repeated clicks, no baseline babysitting. It is what keeps visual testing cheap when the suite has thousands of checkpoints, not three.
Best practice. Work through every difference in the batch first, accepting and rejecting, then save once. Save commits every pending decision in the batch, and it is the one step that is hard to undo, so do it when the counts match what you expect.
๐ Learn more: Test Manager and baselines
Chapter 5: AI closes the loopโ
The idea. Normally this is where the context switch happens: screenshot in one tool, code in another, and you hunting for the CSS rule that only misbehaves at 390px. Instead, the assistant fixes it without leaving the IDE. Through Eyes MCP it can see the diff, the affected element, and the underlying DOM, so it knows where to look.
What you doโ
-
Open the bug. One diff Auto-Maintenance didn't touch: on mobile, the search button is clipped. Not the logo change, a real regression: exactly the kind of diff you would reject and hand off as a ticket. Instead, skip the ticket-to-dev handoff. A bug card appears in the chat; click "Fix the mobile search button."
-
Read the diagnosis. The assistant pulls the diff details for the failing step and reports:
- Root cause: a
max-widthrule on the search form, applied below 400px, was introduced alongside the larger logo and pushed the button past the container edge. - Fix: remove the rule. One line; desktop unaffected.
- Click "Let's rerun and confirm the fix." Run #3 comes back 12 of 12 passed: the logo compares against the baseline you accepted, and the mobile step matches now that the rule is gone. Green you can trust, because you watched Eyes earn it.
โ Takeaway Green across the board: 3 tests, 4 environments, zero regressions. 12 of 12 checkpoints passed. Look at what you just did in a few minutes: replaced 68 assertions with 3 visual checks, covered 4 environments from one run, caught a mobile bug your assertions would never have seen, cleaned up the intended change in one click, and fixed the bug from the chat. That's the loop: test, catch, decide, fix, all in your IDE.
๐ Learn more: Root cause analysis
Take it to your own appโ
The tutorial is simulated so it is free, fast, and safe to replay. Everything it shows is exactly how Eyes behaves in a real repo. Here is the real version, start to finish.
1. Prerequisitesโ
- Node.js 18 or newer and Git
- A Playwright Test project in JavaScript or TypeScript (Eyes MCP setup tools target Playwright today; the Eyes SDKs themselves cover many more frameworks, see below)
- An Applitools account and your API key
2. Add Eyes to your projectโ
Starting from scratch? Clone the ready-made example:
git clone https://github.com/applitools/example-eyes-playwright-fixture.git
cd example-eyes-playwright-fixture
npm install
npx playwright install chromium
Already have a Playwright project? Add the SDK on a branch and review the diff:
npm install -D @applitools/eyes-playwright
npx eyes-setup
npx eyes-setup rewrites your test imports to @applitools/eyes-playwright/fixture and writes the eyesConfig block into playwright.config.ts, so run it on a clean tree and read what changed.
3. Set your API keyโ
export APPLITOOLS_API_KEY='<your key>'
Do this before anything else: a missing key is behind most first-run failures, and the error it produces doesn't say so. Then ask your assistant to "verify my Applitools API key" before you add a single check.
4. Connect Eyes MCP to the assistant you already useโ
Pick the client you already work in; there is no advantage to switching for this.
Claude Code (one line):
claude mcp add applitools-mcp npx @applitools/mcp@latest
VS Code:
code --add-mcp '{"name":"applitools-mcp","command":"npx","args":["-y","@applitools/mcp@latest"]}'
Cursor (~/.cursor/mcp.json), GitHub Copilot (~/.copilot/mcp-config.json), and Cline (cline_mcp_settings.json) all take the same JSON block in their own config file:
{
"mcpServers": {
"applitools-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@applitools/mcp@latest"]
}
}
}
Claude Desktop: add the same server via Settings โ Developer โ Edit Config, per the MCP user guide.
5. Prompts that map to what you just learnedโ
Once the server is connected, these are real prompts, and each maps to a chapter of the tutorial:
| Say to your assistant | What happens | Chapter |
|---|---|---|
| "Verify my Applitools API key" | Validates credentials and connectivity | Setup |
| "Add visual verifications to my existing tests" | Inserts eyes.check() calls following best practices | 1 |
| "Configure the Ultrafast Grid with Chrome, Firefox and Safari at 1280ร800 and iPhone 14 Pro portrait" | Writes the browsersInfo block. Be specific: left to itself the tool proposes its own default environment set | 2 |
| "Summarize the latest visual differences" | Fetches the batch results and reports what changed in plain language | 3 |
| "Review the last batch and just tell me what changed" | Reads and reports without touching any decision | 3โ4 |