Shopify Just Released The Greatest AI Coding Workflow Ever
Discover Shopify's groundbreaking AI coding workflow! Learn how they rebuilt their app with AI, the secrets behind Helix, checkpoints, and gates. Replicate it yourself!
Shopify has recently overhauled its primary mobile application, a significant undertaking that involved rebuilding approximately 300 screens. This extensive project was executed using an AI coding workflow, a structured process that ensured the app's quality and adherence to standards. While Shopify has detailed the methodology behind this workflow, they have not publicly demonstrated its implementation. This article outlines the steps involved in this AI-driven development process, based on a recreated workflow tested on a demo HR system project.
The Helix Workflow: Checkpoints and Gates
Shopify's internal tool for this process is named Helix. The core of the Helix workflow involves breaking down a large task, such as converting a single screen from an old app to a new version, into smaller, manageable units called checkpoints. Before an AI agent proceeds with developing a checkpoint, a human reviewer must approve the proposed tasks. Subsequently, each checkpoint must pass through four rigorous checks, referred to as "gates," before the next checkpoint can commence.
A key element of this workflow is the use of sub-agents. This strategy is employed to manage the context window effectively. When an AI agent handles a large task within a single context window, it can lead to a degradation of work quality as the agent may overlook crucial details. Sub-agents mitigate this by maintaining separate context windows for each specific task, thereby preserving focus and improving output.
The entire loop is fundamentally built around two components: the checkpoints, representing discrete work units, and the four gates, which act as quality assurance checks. Any work failing a gate is halted from progressing to the subsequent task until the gate is cleared. While this workflow was used for app conversion, its principles are applicable to building new features, ensuring adherence to app standards and aesthetic quality.
Recreating the Workflow: The Orchestrator and Checkpoint Planning
Although Helix remains an internal Shopify tool, their released setup has enabled the recreation of a similar workflow. This recreated process was tested on a demo HR system project, focusing on adding features to an existing application. To streamline the execution of this multi-step loop, a skill named "the orchestrator" was developed. The orchestrator acts as a central controller, requiring minimal user input—only at two key stages.
Checkpoints: Deconstructing Features
Checkpoints are the granular components of a larger feature. The workflow mandates that a checkpoint must be fully completed before the next one can begin. Shopify's approach involved ordering checkpoints by increasing complexity, ensuring that smaller, foundational tasks are addressed first. This sequential development allows for early detection and correction of errors, which is more cost-effective when issues are identified at an early stage.
To facilitate this breakdown, a "Checkpoint Planner" agent was created. This agent takes a large feature request and divides it into checkpoints, outputting them in a JSON file format for easy agent comprehension. However, direct human review of the raw JSON file can be challenging due to its structured text format. To address this, a simple webpage viewer was developed to present the planned checkpoints in a more human-readable sequence, indicating progress and task requirements. The Checkpoint Planner is instructed to use simple language for clarity.
Concurrently, another sub-agent is tasked with planning tests for each checkpoint. This documentation ensures that testing strategies are established from the outset. Thorough review of these checkpoint plans is crucial to confirm that no features are omitted and that the subsequent development effort is not wasted.
The Gates: Ensuring Quality at Every Stage
Once a checkpoint plan is approved, the orchestrator initiates the development process, starting with checkpoint one. Before a checkpoint is considered complete, it must pass through a series of review points, or "gates." Unlike simple rules within prompts, gates are mandatory checks that prevent an agent from advancing to the next task until they are satisfied. A hook mechanism was implemented to enforce this, triggering an exit code that prompts the agent to continue until all gates are passed.
Shopify's workflow incorporates four primary gates:
- Behavior Gate: Verifies the functional correctness of the feature.
- UI Gate: Ensures the visual appearance matches the design specifications.
- Code Review Gate: Assesses the quality and integrity of the underlying code.
- Final Review: A human review to confirm user experience and overall functionality.
Gate 1: Behavior Gate
The behavior gate focuses on testing that can be performed without direct browser interaction. Instead of manual clicking, the agent writes code-based tests that simulate user interactions. These tests, planned in advance, establish a clear standard for the agent to verify feature functionality against various inputs. A modified "TDD Planner" skill was adapted to generate and write these tests, covering aspects that do not require a browser, such as permission checks for administrative actions.
When the orchestrator begins checkpoint one, sub-agents are initiated to write the necessary tests for the behavior gate. Initially, these tests will fail as the feature code has not yet been written. Following test generation, another sub-agent writes the actual code for the checkpoint. Finally, a dedicated sub-agent runs these tests, and if all pass, the behavior gate is cleared for that checkpoint.
Gate 2: UI Gate
The UI gate is critical for user experience. In Shopify's app conversion, maintaining design fidelity was paramount. They utilized Gemini models, noted for their spatial awareness, to compare screens and identify subtle differences in spacing and size. For projects without a direct design to compare against, the recreated workflow employs HTML prototypes.
A "prototype" skill generates these prototypes as single, interactive files. These prototypes are built following project design guidelines, often derived from a structured design file. A separate skill then compares the built app against the prototype, ensuring visual consistency and identical behavior upon interaction. This comparison also verifies that both the prototype and the app are in the same state (e.g., a form before and after submission).
For this comparison, the recreated workflow initiates two Claude sub-agents. One agent reviews the screen's visual appearance, and the other assesses its behavior, both judging against the prototype and design file, not the project's instructions. Their combined reviews determine if the UI gate passes. This gate is not applied to every checkpoint, as Gate 1 already covers many functional aspects.
Gate 3: Code Review Gate
Even with functional and visual correctness, the underlying code may require refinement. The third gate, the adversarial review gate, addresses this. Shopify employs two adversarial agents to scrutinize code against established standards, requiring all identified issues to be resolved.
The recreated workflow uses an adversarial loop where one agent identifies issues, and a second "fixer" agent resolves them. A skill named "adversarial loop" manages the communication between these two agents. After the behavior and UI gates are passed, the orchestrator initiates the adversarial agent. If issues are found, the fixer agent addresses them, followed by another review from the adversarial agent. This cycle continues until the adversarial agent approves the code.
Gate 4: Final Review
The final gate involves human oversight. While AI agents operate with defined standards, they cannot fully replicate human perspective. This gate ensures the feature functions as intended from a user's viewpoint and allows for suggestions for improvement.
Once all checkpoints are completed, the user tests the feature. Any requested changes are converted into new checkpoints, which then undergo the entire gate process. Feedback is also recorded in a "learnings file" accessible to all agents before they begin new tasks.
The Orchestrator and Future Development
The entire workflow is managed by the orchestrator skill. By prompting this skill with a task, it guides the process from checkpoint planning to final feature delivery. The orchestrator pauses for user input only after checkpoint planning and upon completion for final review.
Shopify's philosophy is that an "attempt is allowed to be wrong. It is not allowed to ship until it isn't." This structured workflow, with its mandatory gates, ensures that even if an agent makes errors, the final product meets established quality benchmarks. The skills developed for this workflow are available through the AI Labs Pro community.
Introduction to Shopify's AI Coding Workflow
Shopify rebuilt its main mobile app using an AI coding workflow called Helix. This involved breaking down the task into smaller steps and using a structured process with checkpoints and gates to maintain quality, even though the full method wasn't initially revealed. The video aims to demonstrate how to replicate this workflow.
- Shopify rebuilt its 300-screen mobile app using an AI coding workflow.
- The workflow is structured, not a single AI command.
- The method was reverse-engineered and rebuilt by the video creators.
- The workflow uses checkpoints and strict checks ('gates') to ensure quality.
- A custom change was made to fix issues with the AI agent's performance.
Understanding Helix: Checkpoints and Gates
Shopify's Helix tool breaks down app rebuilding into manageable tasks called checkpoints. Each checkpoint undergoes rigorous review through four 'gates' before proceeding. Sub-agents are used to maintain fresh context windows, preventing AI performance degradation on long tasks. This structured approach ensures high standards are met throughout the development process.
- Shopify's tool for the workflow is called Helix.
- Helix converts one screen at a time and splits work into checkpoints.
- Checkpoints are reviewed by a human and must pass four strict checks ('gates').
- Sub-agents are used to manage separate context windows, improving AI performance.
- The workflow relies on checkpoints and four gates for quality assurance.
The Orchestrator and Checkpoint Planning
The video introduces a custom 'orchestrator' skill designed to manage the AI coding loop. This skill requires minimal human input, handling most tasks autonomously. The process begins with a 'Checkpoint Planner' agent that breaks down features into JSON-formatted checkpoints, which are then reviewed by humans via a web viewer. A separate agent plans tests for each checkpoint.
- A custom skill called 'orchestrator' manages the AI coding loop.
- The orchestrator requires human input only twice in the entire process.
- A 'Checkpoint Planner' agent splits features into checkpoints, outputting them in JSON format.
- A web-based viewer is provided for human review of the planned checkpoints.
- A separate agent plans detailed tests for each checkpoint.
Checkpoint Structure and Planning Details
Checkpoints are small, sequential parts of a larger feature, built in increasing order of complexity. This ensures early errors are caught and fixed cheaply. The 'Checkpoint Planner' creates these checkpoints in JSON format for agent readability. A web viewer aids human review, and a separate agent plans tests to ensure thorough validation.
- Checkpoints are small, sequential components of a feature.
- Work is divided into checkpoints of increasing complexity.
- Building smaller tasks first allows for early detection and correction of errors.
- Checkpoints are written in JSON format for easy agent parsing.
- A web viewer and test planning are integral to the checkpoint process.
The Four Gates: Ensuring Quality and Standards
Gates are crucial checks that work must pass before moving to the next stage. Unlike simple rules, gates block progress until cleared. The workflow implements four gates: Behavior (functionality), UI (visual design), Code Review (code quality), and Final Human Review. A custom 'hook' ensures agents cannot proceed until gates are passed.
- Gates are mandatory checks that halt progress if not passed.
- Shopify's workflow uses four gates: Behavior, UI, Code Review, and Final Review.
- A custom 'hook' enforces gate completion before proceeding.
- The Behavior Gate verifies functionality through automated tests.
- The UI Gate ensures the visual design matches the prototype or specifications.
Gate 1: Behavior Gate (Functionality Testing)
The Behavior Gate focuses on testing functionality without browser interaction. Automated tests, planned beforehand, verify the feature's core logic. A modified TDD planner skill generates these tests. The orchestrator then initiates sub-agents to write failing tests, followed by code generation, and finally runs the tests to ensure all pass before marking the gate complete.
- The Behavior Gate tests functionality using code, not manual interaction.
- Automated tests are planned early to provide a clear standard for the AI.
- A modified TDD planner skill generates these tests.
- The orchestrator first generates failing tests, then the code, then runs tests.
- Passing all tests marks the Behavior Gate as complete for a checkpoint.
Gate 2: UI Gate (Visual and Behavioral Consistency)
The UI Gate ensures the visual appearance matches the design. While Shopify used Gemini models for spatial awareness, the replicated workflow uses HTML prototypes. A 'prototype' skill builds these, and a separate comparison skill, using Claude, checks visual and behavioral consistency against the prototype and design file.
- The UI Gate ensures the app's visual design matches requirements.
- Shopify used Gemini models for UI review; the video uses HTML prototypes.
- A 'prototype' skill creates interactive HTML prototypes.
- A comparison skill uses Claude to check visual and behavioral consistency.
- The comparison is made against the prototype and design file, not project instructions.
Gate 3: Code Review Gate (Adversarial Code Analysis)
The Code Review Gate, or adversarial review, scrutinizes the underlying code quality. The video's implementation uses an adversarial loop where one agent identifies issues, and a 'fixer' agent resolves them. This cycle continues until the adversarial agent approves the code, ensuring robustness beyond basic functionality and UI.
- The Code Review Gate checks the quality and robustness of the underlying code.
- Shopify used two adversarial agents for code review.
- The replicated workflow uses an adversarial loop: one agent finds issues, another fixes them.
- The 'adversarial loop' skill manages the interaction between the two agents.
- The gate is passed only when the adversarial agent approves the code.
Gate 4: Human Review (Final User Validation)
The final gate is the Human Review, essential for ensuring the feature works from a user's perspective and identifying potential improvements. After all checkpoints are completed and reviewed, the user tests the feature. Feedback is incorporated as new checkpoints, and learnings are stored in a file for future AI reference.
- The Final Review Gate requires human testing and evaluation.
- Humans assess functionality and user experience from a real-world perspective.
- Feedback can lead to new checkpoints and iterative improvements.
- Learnings from human review are stored in a file for AI agents.
- This gate ensures the feature meets human usability standards.
Conclusion: The Orchestrator and Workflow Summary
The entire AI coding workflow is managed by the 'orchestrator' skill, which automates the process from checkpoint planning to final review. It requires human approval at the checkpoint planning stage and after the final feature is built. The system ensures that while AI attempts can be flawed, no incomplete work is shipped, maintaining high development standards.
- The 'orchestrator' skill manages the entire AI coding workflow.
- Human interaction is limited to checkpoint approval and final feature review.
- The workflow ensures that work must pass all gates before being considered complete.
- Shopify's philosophy: 'an attempt is allowed to be wrong. It is not allowed to ship until it isn't.'
- The workflow can be replicated by prompting Claude or using provided skills.
