Skip to main content

Eyes Interactive Tutorial Guide

Short on time? Everything the interactive tutorial covers lives on this page: every component, every action, and every takeaway. Read it in about ten minutes, come back to it as a reference, or use the last section to run the whole flow on your own app.

The interactive tutorial walks you through your first visual test with Applitools Eyes: you take a Playwright test that checks a whole page in one line, run it on every browser and device at once, and watch it catch a bug that a list of assertions would have missed. This guide covers the same ground in writing.

Contents

  1. Meet your workspace
  2. The flow at a glance
  3. Chapter 1: Write tests faster
  4. Chapter 2: One test, multiple environments
  5. Chapter 3: Visual AI finds the real bug, not the noise
  6. Chapter 4: Auto maintenance
  7. Chapter 5: AI closes the loop
  8. Take it to your own app
  9. Works with what you already use
  10. Best practices recap
  11. Learn more

Meet your workspaceโ€‹

The tutorial simulates a working IDE, so everything you see maps to something you already have. There are five pieces, and only one of them is new.

The five pieces: you prompt an AI assistant inside your IDE, the assistant calls the Eyes MCP server, and the dashboard holds the evidence

PieceWhat it isIn the tutorial
YouYou click, pick, and approve. Nothing runs without you.Buttons and prompts in the chat
Your AI assistantThe coding assistant you already use: Cursor, GitHub Copilot, Claude Code, VS Code, and othersSimulated, so it is free and fast. No real tokens are used
Eyes MCPThe bridge. It gives your assistant Applitools tools: add visual checks, configure browsers, read results. Everything it does can also be done by hand; the MCP removes the manual workInstalled with one command in Chapter 1
Your IDEWhere the code livesEmbedded as two tabs: Website (the demo Wikipedia app and its HTML) and Code Editor (the Playwright test)
Eyes dashboardA separate Applitools web app holding screenshots, baselines, and batch results. It is not part of your IDE: in real life it lives in your browserShown as the Test Results tab: an embedded view of the real Eyes dashboard, pre-loaded for you

One rule of thumb carries the whole tutorial: you talk to the assistant, the assistant uses Eyes MCP, and Eyes MCP does the testing work you would otherwise do by hand. The dashboard is where you look when you want the evidence.

Mock of the tutorial screen: workspace tabs on the left, the assistant chat on the right

A few things to notice as you go:

  • The progress bar ("Your journey") at the top of the chat tracks the five chapters.
  • Tool chips in the chat show every Eyes MCP call by name, so you learn what the server actually does while you watch it work.
  • Locked tabs unlock as the story needs them: the Code Editor opens in Chapter 1, Test Results after your first run.
  • Highlighted elements: when the tutorial pulses part of the screen, clicking it is how you advance. You are clicking the real dashboard, not a picture of it.

The flow at a glanceโ€‹

The loop: check, run, review, decide, save, fix, rerun

The whole tutorial is three runs of the same test, and the story is in the difference between them:

  • Run #1 photographs three pages across four environments and saves those 12 pictures as the baseline: the reference for what "correct" looks like.
  • Run #2 happens after the app changes: one change on purpose, one by accident, and one that only looks like a change. Eyes shows you which is which. You accept the deliberate one and leave the accident flagged.
  • Run #3 proves the accident is fixed and nothing else moved. 12 of 12 pass.

Chapter 1: Write tests fasterโ€‹

The idea. A locator assertion covers exactly the one thing it names. Covering a page properly that way means writing, and then maintaining, dozens of assertions per page, and they still miss everything nobody thought to assert on. One visual check covers the whole page instead: layout, position, size, spacing, color, images, fonts, and how text actually renders, including every element nobody thought to name.

What you doโ€‹

  1. Install Eyes MCP. One command, shown with a copy button:
npx -y @applitools/mcp@latest

The chat confirms the connection, and from here on the โš™ tool chips are how you know Eyes MCP is working.

  1. Open the Code Editor tab and look at the existing Playwright test. The tutorial highlights every await expect(...) line: the assertion list you currently maintain.

  2. Click "Replace the assertions with an Eyes check." The assistant swaps the assertion block for a single call:

import { test } from '@applitools/eyes-playwright/fixture';

test('Wikipedia portal', async ({ page, eyes }) => {
await page.goto(url);

await eyes.check('Wikipedia portal', { fully: true });
});
  1. Click "Add Eyes checks to the other pages." Three pages, three files, and each file is the same shape: go to the page, check the page.

What that one line doesโ€‹

  • fully: true captures the entire scrollable page, so the footer is checked as carefully as the header. No scrolling code.
  • The check name ('Wikipedia portal') is just the label results are filed under.
  • Note what is not in the test: no locators, no selectors, no class names, nothing that breaks when a developer renames something.

โ˜… Takeaway In the demo suite, 68 functional assertions just became 3 visual checks: about 96% less code and locators to write and maintain, for as long as this test exists. That's hours of work saved after every UI change, functional and visual coverage in one check, and less time reviewing test code an LLM generated for you. And every one of those assertions was something a person had to think of in advance; the check catches the things nobody thought of.

Good to know. Nothing here is Playwright-specific. Eyes has SDKs for Playwright, Selenium, Cypress, WebdriverIO, and native mobile, and the same one-line check works in whatever framework your team already uses. There is nothing to rewrite.

๐Ÿ“š Learn more: Applitools MCP server ยท Playwright quickstart


Chapter 2: One test, multiple environmentsโ€‹

The idea. Your users don't all use Chrome on a laptop. Normally, covering Firefox, Safari, and mobile means running the suite again for each one, on a device farm you have to sign up for and maintain. The Ultrafast Grid works differently: it captures your test once, then renders that capture in parallel on every browser, device, OS, and viewport you pick, in the Applitools cloud. No re-running, no grid to host.

One local run fans out to Chrome, Firefox, Safari and iPhone 14 Pro: 3 checks times 4 environments is 12 checkpoints

How the Ultrafast Grid works. When your test runs on your machine, in one browser, Eyes doesn't just take a screenshot. It captures the page itself: the DOM, the CSS, and the resources it loads. The Ultrafast Grid then takes that capture and renders it on real browser engines in the Applitools cloud, in parallel, on every environment you configured: desktop browsers (Chrome, Firefox, Safari, Edge) at the viewport size you set, and mobile devices at their native screen sizes in portrait or landscape. Because the navigation runs once, and not on each device, the process is significantly faster and more stable than re-running the test per environment, and there is no device farm, Selenium grid, or QA lab to maintain. Adding an environment is one line of config, not another test run.

What you doโ€‹

  1. Pick your environments. The chat shows a browser picker. Chrome and iPhone 14 Pro are always included; Firefox and Safari are one click each. Your selection becomes a browsersInfo block in playwright.config.ts:
eyesConfig: {
appName: 'Wikipedia demo',
type: 'ufg',
browsersInfo: [
{ name: 'chrome', width: 1280, height: 800 },
{ name: 'firefox', width: 1280, height: 800 },
{ name: 'safari', width: 1280, height: 800 },
{ iosDeviceInfo: { deviceName: 'iPhone 14 Pro', screenOrientation: 'portrait' } },
],
}

That block is the only place the environment list lives. One config, four environments, still one test.

  1. Click "Run the test." Because this is the first run, Eyes has nothing to compare against yet, so it saves what it sees as the baseline: the reference picture of "correct", per checkpoint, per environment. Everything comes back Passed, and that is expected.

  2. Open the batch in the dashboard. The result card links straight to it. Confirm three things: 4 environment rows, 12 checkpoints, all Passed. Click into the iPhone 14 Pro row: that is a real mobile rendering of a capture taken once.

Vocabulary the dashboard usesโ€‹

  • A batch is one run of your suite.
  • Each row is a test in a specific environment.
  • Each cell is a checkpoint: the screenshot Eyes captured for one check.
  • A baseline is the reference image a checkpoint is compared against. Baselines live in your Applitools account, not in your repo.

โ˜… Takeaway One test. 4 environments. 12 baselines. Eyes rendered your single run everywhere at once: no re-running, no device cloud sign-up, no grid to touch. From now on, every run checks every environment.

  • About 10ร— faster than running per browser, so CI stops waiting on cross-browser coverage
  • No secure tunnels to punch through your firewall
  • No timeout or concurrency handling to write: name the devices and browsers, the SDK does the rest

Best practices.

  • With the Ultrafast Grid, keep a single Playwright browser project. If you keep multiple local projects, they multiply by grid environments and every result appears twice.
  • Treat the first run as a deliberate act: open the batch and look at what was captured before you accept it as the definition of correct.

๐Ÿ“š Learn more: Ultrafast Grid ยท Configuration options ยท Batching


Chapter 3: Visual AI finds the real bug, not the noiseโ€‹

The idea. Pixel diffing finds everything that changed. Visual AI finds what matters. This chapter is the heart of the tutorial.

What you doโ€‹

  1. A design request lands. A ticket card appears in the chat: "Increase the Wikipedia logo by 1.2ร— on all pages." In a real project this could come from Jira, Slack, or a teammate; the flow is the same. Click "Make this change" and the assistant edits the app's CSS. A one-line change any of us would ship on a Friday, and nobody is going to re-test mobile for a logo resize.

  2. Click "Rerun the test on the same environments." Same 3 checks, same 4 environments, now compared against the baselines from Run #1. The result: 9 of 12 checkpoints unresolved. Unresolved is Eyes' honest third status: a visual difference that needs a human decision, bug or intentional change.

  3. Click "Inspect the batch result." Instead of clicking through 12 screenshots, the assistant reads the batch through Eyes MCP and reports back in plain language:

  • Real change: the logo is about 1.2ร— larger on every page that shares the stylesheet. 8 steps, same underlying change everywhere. Not 8 separate issues.
  • Possible regression: on iPhone 14 Pro only, the Wikipedia portal's search button is clipped. It is not part of the design request.
  1. Open the (view) link and switch to the pixel view (Exact). The whole page lights up: the logo, plus every line of navigation and body text that shifted by half a pixel when the logo grew. This is what any screenshot-diff tool would report. Dozens of failures for one intended change. This is the noise that makes teams stop trusting visual tests.

  2. Click "Back to Visual AI (Strict) view." Same page, same pixels. One highlight instead of dozens: the logo. The half-pixel shift is gone from the report because it isn't a defect, and the clipped mobile button is still flagged because it is.

Pixel diff lights up the whole page; Visual AI highlights the one real change and still flags the clipped mobile button

Match levels: how picky should a check be?โ€‹

A match level is the rule Eyes uses to decide what counts as a difference. It is a per-check decision, not a global setting:

LevelFlagsReach for it when
Dynamic (default for new tests)Like Strict, but auto-recognizes dynamic content (dates, IDs, counters, currency) and ignores the displacement it causesAlmost always. Real pages have content that changes on every load
StrictAnything a person would notice: text, color, layout, sizeBrand, legal, forms; anywhere exact rendering is the point
LayoutStructure only. Content changes are ignored, but a vanished, overlapping, or resized element is still flaggedSearch results, feeds, anything data-driven
Ignore ColorsContent and layout; color-only drift is toleratedMid-rebrand, themed surfaces

Set it per check in code:

await eyes.check('Search results', { fully: true, matchLevel: 'Layout' });

โ˜… Takeaway Pixel diff, like other tools, found everything that changed. Applitools Visual AI found what matters: the change you asked for, and a real bug you didn't. Fewer false alarms means every flag is worth opening, and match levels like Strict, Layout and Dynamic let you tune that line per check. And it is deterministic: same input, same result, every run. No guessing.

Best practice. Start on Dynamic for a real application. Tighten to Strict where exact rendering matters. Loosen a region to Layout when its content is data-driven but its structure isn't. Reach for an Ignore region last: it is the only option that makes you blind to that area entirely.

๐Ÿ“š Learn more: Match levels ยท Hide displacement diffs


Chapter 4: Auto maintenanceโ€‹

The idea. Low-maintenance testing that scales with the team. The logo change is intended, so the new screenshots should become the new baseline. With most tools you would now accept it on every page and every environment, one by one: 8 times here, hundreds of times on a real suite.

What you doโ€‹

  1. Accept it once. The tutorial highlights the Accept button on one logo checkpoint, in the real dashboard. Click it.

  2. Watch what Eyes does with the others. The other logo checkpoints flip to Accepted one after another, each tagged Auto-maintained. Eyes recognized the same difference on the other pages and environments and applied your decision there too.

  3. Notice the one that stays unresolved. The Wikipedia portal checkpoint on iPhone doesn't flip, because its diff contains something you haven't ruled on: the clipped search button. That is the difference between Auto-Maintenance and a bulk-accept button. Your decision propagates only where the difference matches, and is withheld where it doesn't.

Accept one checkpoint and seven auto-maintain; the one with a different diff stays unresolved

  1. Save the batch. The tutorial highlights the Save button. Saving is what turns decisions into baselines: accepted checkpoints become the reference for the next run, and anything left flagged keeps failing until the code changes. Until you save, nothing has moved.

โ˜… Takeaway One click, 7 more resolved on their own. That's Auto-Maintenance: no repeated clicks, no baseline babysitting. It is what keeps visual testing cheap when the suite has thousands of checkpoints, not three.

Best practice. Work through every difference in the batch first, accepting and rejecting, then save once. Save commits every pending decision in the batch, and it is the one step that is hard to undo, so do it when the counts match what you expect.

๐Ÿ“š Learn more: Test Manager and baselines


Chapter 5: AI closes the loopโ€‹

The idea. Normally this is where the context switch happens: screenshot in one tool, code in another, and you hunting for the CSS rule that only misbehaves at 390px. Instead, the assistant fixes it without leaving the IDE. Through Eyes MCP it can see the diff, the affected element, and the underlying DOM, so it knows where to look.

What you doโ€‹

  1. Open the bug. One diff Auto-Maintenance didn't touch: on mobile, the search button is clipped. Not the logo change, a real regression: exactly the kind of diff you would reject and hand off as a ticket. Instead, skip the ticket-to-dev handoff. A bug card appears in the chat; click "Fix the mobile search button."

  2. Read the diagnosis. The assistant pulls the diff details for the failing step and reports:

  • Root cause: a max-width rule on the search form, applied below 400px, was introduced alongside the larger logo and pushed the button past the container edge.
  • Fix: remove the rule. One line; desktop unaffected.
  1. Click "Let's rerun and confirm the fix." Run #3 comes back 12 of 12 passed: the logo compares against the baseline you accepted, and the mobile step matches now that the rule is gone. Green you can trust, because you watched Eyes earn it.

โ˜… Takeaway Green across the board: 3 tests, 4 environments, zero regressions. 12 of 12 checkpoints passed. Look at what you just did in a few minutes: replaced 68 assertions with 3 visual checks, covered 4 environments from one run, caught a mobile bug your assertions would never have seen, cleaned up the intended change in one click, and fixed the bug from the chat. That's the loop: test, catch, decide, fix, all in your IDE.

๐Ÿ“š Learn more: Root cause analysis


Take it to your own appโ€‹

The tutorial is simulated so it is free, fast, and safe to replay. Everything it shows is exactly how Eyes behaves in a real repo. Here is the real version, start to finish.

1. Prerequisitesโ€‹

  • Node.js 18 or newer and Git
  • A Playwright Test project in JavaScript or TypeScript (Eyes MCP setup tools target Playwright today; the Eyes SDKs themselves cover many more frameworks, see below)
  • An Applitools account and your API key

2. Add Eyes to your projectโ€‹

Starting from scratch? Clone the ready-made example:

git clone https://github.com/applitools/example-eyes-playwright-fixture.git
cd example-eyes-playwright-fixture
npm install
npx playwright install chromium

Already have a Playwright project? Add the SDK on a branch and review the diff:

npm install -D @applitools/eyes-playwright
npx eyes-setup

npx eyes-setup rewrites your test imports to @applitools/eyes-playwright/fixture and writes the eyesConfig block into playwright.config.ts, so run it on a clean tree and read what changed.

3. Set your API keyโ€‹

export APPLITOOLS_API_KEY='<your key>'

Do this before anything else: a missing key is behind most first-run failures, and the error it produces doesn't say so. Then ask your assistant to "verify my Applitools API key" before you add a single check.

4. Connect Eyes MCP to the assistant you already useโ€‹

Pick the client you already work in; there is no advantage to switching for this.

Claude Code (one line):

claude mcp add applitools-mcp npx @applitools/mcp@latest

VS Code:

code --add-mcp '{"name":"applitools-mcp","command":"npx","args":["-y","@applitools/mcp@latest"]}'

Cursor (~/.cursor/mcp.json), GitHub Copilot (~/.copilot/mcp-config.json), and Cline (cline_mcp_settings.json) all take the same JSON block in their own config file:

{
"mcpServers": {
"applitools-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@applitools/mcp@latest"]
}
}
}

Claude Desktop: add the same server via Settings โ†’ Developer โ†’ Edit Config, per the MCP user guide.

5. Prompts that map to what you just learnedโ€‹

Once the server is connected, these are real prompts, and each maps to a chapter of the tutorial:

Say to your assistantWhat happensChapter
"Verify my Applitools API key"Validates credentials and connectivitySetup
"Add visual verifications to my existing tests"Inserts eyes.check() calls following best practices1
"Configure the Ultrafast Grid with Chrome, Firefox and Safari at 1280ร—800 and iPhone 14 Pro portrait"Writes the browsersInfo block. Be specific: left to itself the tool proposes its own default environment set2
"Summarize the latest visual differences"Fetches the batch results and reports what changed in plain language3
"Review the last batch and just tell me what changed"Reads and reports without touching any decision3โ€“4

Works with what you already useโ€‹

Eyes SDKs. The tutorial uses Playwright, but the same one-line check works across the stack your team already has: Playwright, Selenium, Cypress, WebdriverIO, and native mobile, across JavaScript/TypeScript, Java, Python, C#, Ruby, and more. Results from any SDK land in the same dashboard, with the same baselines, Auto-Maintenance, and review flow. See the quickstarts for your framework.

Editors and AI assistants. Eyes MCP speaks the Model Context Protocol, so it plugs into the assistant you already use:

  • VS Code + GitHub Copilot
  • Cursor
  • Claude Code (terminal, works alongside any editor)
  • Claude Desktop
  • Cline
  • Any other MCP-capable client, using the standard JSON config above

Requirements for the MCP server: Node.js 18+, and a Playwright JS/TS project for the setup tools.


Best practices recapโ€‹

Everything the tutorial teaches by doing, in one list:

  1. Let one check cover the page. Prefer eyes.check(name, { fully: true }) over element-by-element assertions. Coverage should not scale with lines of code.
  2. Keep one local browser project and let the Ultrafast Grid do the fan-out. The environment list lives in one browsersInfo block, nowhere else.
  3. Treat the first run as a deliberate act. Look at your baselines before you accept them as the definition of correct.
  4. Unresolved is a feature. Pass/fail is not enough for visual changes; a diff needs a human decision, and Eyes keeps it explicitly pending until you make one.
  5. Pick match levels in order. Dynamic by default, Strict where exact rendering is the point, Layout for data-driven regions, Ignore regions last.
  6. Accept once, let Auto-Maintenance propagate. And trust the one that stays unresolved: it stayed for a reason.
  7. Decide everything, then save once. Saving commits every pending decision in the batch to baseline, so make it the last step of a review, not a reflex.
  8. Fix from the evidence. The diff, the affected element, and the DOM are all available to your assistant through Eyes MCP; a rerun is the proof the fix landed.
  9. Name batches meaningfully and group parallel CI jobs with one batchId so a run shows up as one entry in the Test Manager.

Learn moreโ€‹

Docs

Example projects