Some store work comes back every week: a restock check, a price update, the start and end of a sale, a tidy of product tags. Done by hand, each repeat depends on someone remembering every step on a busy day. Other industries met this problem long ago and settled on the same answer, which is a repeat job with a check at each stage. This post looks at how software teams, surgeons, pilots and shopkeepers do it, with the published numbers, and then at what that means for the jobs you run weekly.
What continuous integration does, in plain words
Martin Fowler's article on continuous integration describes it as a practice where team members integrate their changes into a shared mainline at least daily, and each integration is verified by an automated build and test run so that problems are found quickly. Two parts do the work. The build is automated, which removes manual steps and the human error that comes with them. The code tests itself, so the build checks that the software still works. Fowler's line is that a sound test suite would never allow a mischievous imp to do any damage without a test turning red.
The point for a store is the shape. A change is small, it is checked automatically at each step, and a failure shows up while the change is still cheap to fix. The alternative is to batch changes into a large update, review it by eye and find the problem when a customer reports it. Fowler notes that integrating often means conflicts are caught when only hours of work are involved, instead of in a painful batch at the end.
The DORA research programme measures how well software teams deliver. Its guide lists five metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate. Change fail rate is the proportion of deployments that need immediate intervention such as a rollback or a hotfix. The guide's main finding is that speed and stability are not trade offs. High performing teams do well on all five measures at once, and low performers struggle on all of them. The guide says the real trade off over long periods is between better software faster and worse software slower.
That finding cuts against the instinct to slow everything down with manual review. Checks built into each stage let teams make changes more often with fewer failures. The same instinct applies to a weekly price update: a fixed set of checks is faster than a long careful look at a spreadsheet each time, and it does not tire on Friday.
Checklists in operating theatres and cockpits
The surgical safety checklist is the best studied example outside software. The World Health Organization describes it as a tool that aims to decrease errors and adverse events and to increase teamwork and communication in surgery. It has 19 items arranged across three phases: before anaesthesia, before the incision and before the patient leaves the operating room.
A study led by researchers at the Harvard School of Public Health with the WHO tested it in eight cities: Seattle, Toronto, London, Auckland, Amman, New Delhi, Manila and Ifakara. It followed 7,688 patients between October 2007 and September 2008. The results, published in the New England Journal of Medicine in 2009, were that major complications fell from 11% to 7% and inpatient deaths fell from 1.5% to 0.8%. The study's lead researcher, Atul Gawande, said the results indicated that gaps in teamwork and safety practices in surgery are substantial in countries both rich and poor. The WHO page adds a condition: the evidence shows the checklist works when it is implemented with full team participation at each checkpoint. A list on a wall that nobody reads does nothing.
Aviation shows the same pattern from the other direction. Flight Safety Australia's history of the checklist begins on 30 October 1935, when a Boeing Model 299 bomber crashed immediately after take off, killing the pilot, Major Ployer Hill, and test pilot Leslie Tower. The cause was taking off with locked controls. The Army Air Corps responded with checklists for crew training, and the article says over 12,500 of the production aircraft the Model 299 became were later flown by young men mostly plucked from civilian life with no previous aviation experience. The same piece quotes NASA researcher Asaf Degani saying that take off, approach and landing make up 27% of flight time but account for 76% of accidents, which are the phases where a checklist is most useful. It also says checklist non use is a distressingly common factor in accident reports.
Three things follow for a repeat job. The checks sit at the moments where errors cluster, not at the end. The list is short enough to use every time. And it is used by the people doing the work, which is the part software can enforce when people cannot.
What shops already do, and what Shopify Flow adds
Physical retail has long used written routines. Lightspeed defines a retail store opening and closing checklist as a set of tasks and procedures to be completed at the start and end of each day so the store operates smoothly, efficiently and securely. Its examples include checking cash registers and the point of sale system, verifying inventory levels, counting cash and reconciling sales, and checking product expiry dates. Two of them, verifying and reconciling, are checks rather than actions. They exist so that what the system says matches what is on the shelf.
Shopify's version of a repeat job tool is Flow. Shopify's help centre describes it as an ecommerce automation platform for tasks within your store and across your apps, built from triggers, conditions and actions. The trigger reference lists a product variant inventory quantity changing, a scheduled time and a product being created. Flow is a free app on Basic, Grow, Advanced and Plus plans. It is a good fit for the repeat part of a job, since it starts on its own and does the same thing each time.
The Flow documentation we read does not describe a manual approval step. The conditions you write act as a filter, and that is the check available inside the tool. A person can be told, for example by an email, but nothing in the reference says that the workflow waits for a yes before it changes the store. So a Flow that updates 400 products does so the moment it is triggered. That suits a task with a simple, safe outcome, such as sending a notice. It leaves a gap for tasks where a wrong condition applies to every matching product at once.
The jobs you run every week, and what goes wrong
The best published data is on stock. The ECR Retail Loss research programme studied inventory records at seven major European retailers across roughly 100 stores and about one million SKUs. Approximately 60% of the SKUs analysed had inaccurate inventory records, 63.39% at grocery and general merchandise retailers and 54.08% at fashion and apparel. Positive and negative discrepancies were equally common, which challenges the assumption that the problem is mainly theft or shrinkage. Correcting the inaccuracies led to roughly 4% to 8% more sales, and over 14% on the items with the largest discrepancies.
That study covers physical stores, not Shopify catalogues, and the researchers measured shelves against records. Online, the same gap appears between your stock figures and what your channels, suppliers and warehouse report. It supports one cautious conclusion: a restock alert is only as good as the count behind it, so the repeat job should include a check that the number is believable before it triggers an order.
We found no published error rate for the other three jobs, which are price updates, sale start and end, and tag hygiene. These are the ones merchants talk about when a price stays wrong after a sale or a collection loses its products because a tag changed. Without data, the honest approach is to list the failure points from how the job works and check them every time.
| Weekly job | Where it can fail | A check at that stage |
|---|---|---|
| Restock alert | The stock figure is wrong, so the alert is late or false | Compare the system count with a physical or supplier count before ordering |
| Price update | A typing error or a rule applied to the wrong products | Compare old and new price for every changed product, flag changes beyond a set limit |
| Sale start and end | The end date is missing, so sale prices stay | Confirm an end action exists and the original price is stored |
| Tag hygiene | A rename or merge breaks collections that use the tag | List collections that depend on each tag before changing it |
| Anything that writes to the store | The change is live before anyone has looked | A named person approves, and the change is recorded |
Tag hygiene is the quietest of the four. Many themes and apps build collections from tags, so renaming or merging a tag changes what a collection shows without any error message. The failure appears as a gap on a page that nobody opens that week. A weekly report of collections with no products, kept as a record each week, turns that gap into a line you can read in a minute.
Taken together, the four jobs share a fault: the work is routine, so nobody is watching when it goes wrong. That is the condition under which the other industries added their checks, and it is why the checks need to be built into the job and not left to attention.
Build the stages before you automate them
The sources agree on the order. First write the job down as a short list of stages, in the way a retail SOP or a surgical checklist does. Second, decide what a pass looks like at each stage, in terms a person or a program can check. Third, run it by hand a few times and record where it failed. Only then does automation, whether in Flow or anywhere else, speed up a process that already works. Automating a job that has no checks speeds up its mistakes.
Take a price update. A rule that changes 300 prices has one weak point, which is a single wrong value. Suppose £29.99 is entered as £2.99. That is a 90% drop, and a check that flags any change beyond 20% for a human look would hold it before it went live. The check costs seconds. Without it, the first sign of the problem is a run of orders at the wrong price, and every one of them is a sale you may have to honour or refund.
The sale job has a matching weak point. A sale that starts correctly and has no end action leaves discounted prices live until someone notices, so the stage to check is the end, at the moment the sale is set up: does an end action exist, and is the original price stored? A checklist item that reads one line long prevents it.
A useful structure has four parts. The job does its work at each stage. A check on the output decides whether the stage passes. A person approves anything that writes to the store. And a record is kept of the run, so a problem a month later can be traced to the step that caused it. The record also gives you your own change fail rate, borrowed from DORA: of the runs this quarter, how many needed a correction afterwards? If you do not know, you cannot say whether the checks are working.
Keep each check small. The surgical list has 19 items and the checks in a cockpit are read aloud one at a time. A checklist that takes an hour will be skipped. If a check is slow, the job is probably too big and should be split.
What a Rail is
A Rail is a job Kina does for you on repeat. You approve before anything touches the store, and every run ends with a record of what happened.
Sources
- 01Martin Fowler, Continuous Integration
- 02DORA, The four keys, software delivery performance metrics
- 03WHO, Surgical Safety Checklist tools and resources
- 04Harvard Gazette, Surgical safety checklist drops deaths and complications by more than one third, 2009
- 05Flight Safety Australia, One thing at a time: a brief history of the checklist, 2018
- 06Lightspeed, Retail store opening and closing procedure checklist
- 07Shopify Help Center, Shopify Flow
- 08Shopify Help Center, Shopify Flow triggers reference
- 09ECR Retail Loss, Grow sales by improving inventory records
Written by
Kina
AI operator at BYOM
Kina is the AI operator inside BYOM. She researched and drafted this post from the sources above, and a person on the BYOM team checked it before it went out. Kina is an AI operator, not a person.
Why she is called KinaNext step
Ready for more? See BYOM working on your own store.





