The same operations toolkit David used inside real plants — value stream maps, maintenance calendars, 5S boards, safety audits — now sized for your business. This is how we find what's costing you money before we touch a single tool.
// The operations toolkit
The same tools David used to run $500M+ plants, sized for your business.
Twenty years finding waste on plant floors comes with a real toolkit, not a slogan. Here are nineteen of those tools, with the full breakdown of each one — free, open, right on this page.
Standard, industry-level frameworks — not the specific way we apply them inside a paid engagement. No invented numbers; sample fields are marked as illustrative.
// Tool 01
Value Stream Map (VSM)
See how much of your lead time is real work — and how much is waiting.
Receive → Prepare → Deliver · C/T 10 / 25 / 10 min · wait 3 days vs 45 min of work
Most businesses don't know how much of their delivery time is actual work versus waiting, moving, or sitting in a queue. A VSM draws the full flow, from the customer's order to delivery, and separates what adds value from what doesn't — so improvement dollars go to the real bottleneck, not where the problem "feels" like it is.
Customer/demand box: units per period, order frequency, takt-time calculation
Supplier box: vendor, delivery frequency, lot size
One box per process step: cycle time, changeover time, availability %, operators, shifts
Inventory triangles between steps: quantity and days of coverage
Information flow arrows (manual vs. electronic): orders, scheduling, forecast
Timeline ladder: value-added time vs. wait time per stage, and totals
Kaizen bursts marking each improvement opportunity found
// Tool 02
Autonomous Maintenance (CILR) Calendar
Most "sudden" breakdowns had a warning days earlier.
EQ-04 daily CILR · loose guard bolt on day 9 (R-cycle) · closed day 11
A small leak, a loose screw, a different sound — most equipment failures that look sudden already gave a sign. This calendar gives the person running the equipment, not just the maintenance tech, a daily CILR routine: clean, inspect, lubricate, and refasten — and makes sure it actually happens instead of "when there's time."
Equipment data: ID, area, responsible operator
CILR task list by component: clean / inspect / lubricate / refasten, with frequency, time, and tool needed
Reference point on the equipment (photo or simple diagram)
Check matrix: days of the month in columns, operator initials, OK/anomaly mark per day
Anomaly log: date, description, who was notified, resolution date
Escalation column: operator resolves vs. needs specialized maintenance
// Tool 03
5S Audit & Sustainment Board
5S doesn't fail at the cleanup day. It fails three months later.
Warehouse A · owner M. Ruiz · Shine dip in week 3 · torque wrench missing from the shadow board
5S almost always starts strong with one big cleanup day, then slides back because nobody owns keeping it. This board turns 5S from a one-time event into a management habit — a weekly score per area, with a visible trend and a named owner — so a decaying area gets caught before it's dirty again.
Area/zone, audit date, auditor
The 5 categories (Sort, Set in order, Shine, Standardize, Sustain), each with 3-5 concrete checkpoints, scored yes/no or 0-5
Score per category and total/percentage
Photo evidence field (current state)
Comments, corrective action, owner, commitment date
Trend row: score by week/month, visible on the physical or digital board
// Tool 04
Safety & Environmental Risk Matrix
Find out what your biggest risk is before it finds you.
A chemical storage 4×5 high · B handling 3×3 moderate · C panel 2×2 low after LOTO
Without a structured review, an owner usually finds out what their worst operational risk is after it already happened — an injury, a fine, a spill, a shutdown. This matrix walks every task or area, scores likelihood and severity, and points a limited safety budget at the two or three risks that actually matter.
Task/area/process evaluated
Hazard identified, by category: mechanical, chemical/environmental, ergonomic, electrical, fire
Who or what is exposed: operator, community, environment, equipment
Existing controls, and likelihood / severity scored 1-5 each (severity includes environmental impact, from contained on-site spill to reportable release)
Risk score (likelihood × severity) and color-coded risk level (low/medium/high/critical)
Recommended additional controls, following the hierarchy: eliminate, substitute, engineering control, administrative control, PPE
Owner, target date, residual risk after controls
// Tool 05
OEE Dashboard
Equipment that's "on" isn't the same as equipment that's producing.
PKG-01 · A 88% × P 91% × Q 97% = 78% OEE · changeover 142 min is the loss
Equipment can run without ever fully stopping and still lose a large share of its real capacity to short stops, slow cycles, and defects, unnoticed. This tool turns raw shift data into one number, Availability × Performance × Quality, broken down by cause, so the decision to fix what's on the floor or invest in new equipment is made with data.
Shift/date, equipment, planned production time
Stop log: category (breakdown, changeover, minor stop, startup) and minutes per event
Total parts produced, ideal cycle time, good vs. defective parts
Calculated fields: Availability %, Performance %, Quality %, OEE % (the product of the three)
Pareto of stop causes, ranked by minutes lost in the period
Trend row: OEE by shift/day/week
// Tool 06
Kanban / Pull Board
Stop guessing what to make next — let what's actually being used tell you.
PN-18 · Cut → Assemble → Pack · 10 cards · empty bin at Pack releases work
Most shops build to a forecast, then discover half of what they made is sitting in a corner while a different item runs out. A kanban board replaces that guess with a signal: nothing gets made or moved until the step downstream actually consumes it. It caps how much can pile up between any two steps and makes overproduction visible the moment it starts, instead of three weeks later when it's tying up cash on a shelf.
Item/part number, and the two process steps it flows between (produced-by / consumed-by)
Average demand per period, and its source (order history window, forecast — dated)
Replenishment lead time for that item (measured, not assumed)
Container/lot quantity, and safety factor applied
Calculated card count: (avg. demand × lead time × safety factor) / container quantity
Trigger point: inventory level or empty-bin signal that releases a card
Review cadence: when card count gets recalculated as demand or lead time changes
// Tool 07
KPI Tree
The number on your P&L moved. This is how you find out why, before next month.
Revenue $84k · attendance / conversion / avg ticket · F&B attach is the lever · safety locked
A single result number — revenue, margin, satisfaction — tells you something is wrong but never what to do about it Monday morning. A KPI tree connects that result to two or three process indicators a manager can actually move that same week, and shows the causal link between them instead of a wall of metrics nobody owns. It works the same way for a service business or a retail floor as it does for a production line — the “product” just changes from units to visits, tickets, or bookings.
Top-level (lagging) KPI: what it is, current baseline, and target
2-4 process (leading) KPIs beneath it that a supervisor can move within the week
For each KPI: owner (one named person, not a department)
Review frequency per KPI (daily/weekly/monthly) and where it gets reviewed
Data source for each number, and how it's calculated
Evidence the leading indicator actually predicts the lagging one (historical check, not assumption)
Safety/non-negotiable indicators flagged separately, never traded off against commercial ones
// Tool 08
Fishbone Diagram (Ishikawa)
Before you fix anything, get every possible cause on the table — not just the first one someone blames.
18 late in pack bay · 3 causes verified · priority rule still unwritten
When something goes wrong, the fastest explanation people reach for is usually “someone made a mistake” — and that's rarely the whole story. A fishbone diagram forces a structured look across every category that could be causing the effect (method, machine, material, people, measurement, environment) before anyone commits to a fix. It's a hypothesis map, not a verdict — each branch still needs to be checked against real data before you spend money acting on it.
The effect/problem statement, defined precisely (what, where, when, how much)
Six cause categories: Method, Machine, Material, Manpower (people), Measurement, Environment — or the four adapted for service work: Policies, Procedures, People, Plant/Place
Candidate causes listed under each category, from a group session, not one person's opinion
For each candidate cause, whether it's verified against real data or still a hypothesis
The 2-3 causes selected for deeper analysis (feeds directly into 5 Whys or a Pareto)
Date and team who built it, so it can be revisited instead of redone from scratch
// Tool 09
5 Whys Root Cause Worksheet
If your last five whys all end at “the employee should have been more careful,” you didn't find the cause — you found someone to blame.
Late at 16:12 · unwritten cutoff rule · action is a posted standard, not blame
Asking “why” once gets you a symptom. Asking it five times, each answer checked against what actually happened on the floor rather than assumed from a desk, gets you to something a policy, a design, or a maintenance decision actually caused — and that's the level where a fix sticks. The worksheet forces each step to cite evidence, not opinion, and it's a red flag if every chain ends the same convenient way, at the person doing the work.
Problem statement: specific, with date, location, and measured impact
Why #1 through why #5 (or however many it takes), each one a cause, not a restatement of the problem
Evidence source for each “why”: what was checked on the floor, who confirmed it, when
Point where the chain stops being about an event and starts being about a system (design, policy, standard)
Root cause selected for action, and who validated it wasn't just the first plausible answer
Corrective action tied explicitly to the root cause found, with owner and date
// Tool 10
Pareto Chart Worksheet
Most operations have one or two causes doing most of the damage — find them before you spread the budget thin.
Jul 2026 · 1,012 min down · top 2 = 66% · cut line after printer jams
Fixing everything a little bit usually means fixing the thing that actually matters not at all. A Pareto chart ranks every cause of a problem — downtime, defects, complaints, delays — by how much it actually costs in time or money, not by how loud or recent it was, and shows what share of the total the top few causes account for. Most operations find that two or three causes explain the majority of the loss, which is exactly where a limited budget and a limited team should go first.
Problem/loss category being analyzed (downtime, defects, complaints, safety near-misses, etc.)
List of causes or categories, one row each
Frequency or cost per cause, with the unit and the time window it covers
Causes sorted descending by frequency/cost
Cumulative percentage column, running total across sorted causes
The cut line: which causes make up roughly 80% of the total, marked visually
Action owner assigned to the top causes only, not the whole list
// Tool 11
A3 Problem-Solving Sheet
One page that forces the thinking — not a report written after the decision was already made.
5-day ship vs 2-day promise · gap 3 days · 3 countermeasures · check 14 Oct
Most “root cause” write-ups get built backward: the fix is already decided, and the analysis gets typed up afterward to justify it. An A3 is built the other way — one page that walks context, current state, target, root cause analysis, countermeasures, plan, and a follow-up check, in that order, so a reader can see exactly where the reasoning happened and disagree with a specific step instead of the whole conclusion. The single sheet is a side effect; the real product is the thinking that had to happen to fill it in honestly.
Background: why this problem matters now, and which business objective it connects to
Current state: the process today, with real data — not the version that makes the team look good
Target: specific, measurable, dated
Root cause analysis: the reasoning shown (fishbone, 5 whys, Pareto), not just a stated conclusion
Countermeasures, each one explicitly linked to a specific root cause, not a general fix
Implementation plan: who, what, by when, broken into verifiable actions
Follow-up: how and when the result gets checked again to confirm it held — not just that it was implemented
// Tool 12
Andon / Visual Workplace Board
A problem anyone can flag and everyone can see gets fixed today, not next month.
Week of 11 Aug · printer jam ×3 escalated · cutoff miss Fri 16:12
In most shops, a problem on the floor stays a private frustration until it becomes a crisis — the person who spotted it has no fast way to raise it, so it waits. An andon system gives anyone the ability to signal an abnormal condition the moment it happens — a light, a card, a button — and makes that signal visible to the whole area, not just to whoever happens to walk by. The board only works as long as every signal gets a response logged against it; an andon light nobody answers stops being a tool and becomes decoration.
Trigger types: what conditions call for a signal (quality defect, safety concern, material shortage, equipment issue)
Signal method: visual/audible mechanism used (light, card, button, cord) and its location
Response protocol: who responds, in what time window, and what they're authorized to decide on the spot
Escalation path: what happens if the first responder can't resolve it within the time window
Log: date/time raised, cause, who responded, time to resolution
Recurrence flag: same cause triggering repeatedly, marked for a real fix instead of another quick response
Review cadence: how often the log gets reviewed for patterns, and by whom
// Tool 13
Waste Walk (7 Wastes)
Walk the same floor you walk every day — but this time, count what's actually being wasted.
Pick→pack walk · 6 finds · motion 18 min/shift is #1
Waste hides in plain sight because everyone on the floor has stopped seeing it — the extra trip to fetch a part, the stack of work-in-process waiting between two steps, feels normal because it's always been there. A waste walk is a structured pass through the process, category by category — overproduction, waiting, transport, overprocessing, inventory, motion, defects, and unused talent — that names each instance on the spot instead of relying on memory back in an office. It turns a vague sense that “we could be more efficient” into a specific, countable list ranked by what's actually costing the most.
Area/process walked, date, who walked it (ideally cross-functional, not one person alone)
One row per waste category (the 7 wastes, plus unused talent as the 8th)
Specific instance observed per category: what, where, how often
Estimated impact: time, distance, or units affected, with the basis for the estimate noted
Photo or sketch evidence for each instance, where practical
Priority ranking across all instances found, by estimated impact
Owner and next step assigned to the top few, not the entire list
// Tool 14
Poka-Yoke Design Worksheet
The best fix for human error isn't a warning sign — it's making the error physically impossible.
Wrong insert 6× in July · prevention chosen: keyed tote + WMS lock
Telling someone to “be more careful” doesn't survive a busy shift, a new hire, or hour ten of a long day — the same mistake comes back. Poka-yoke works through error-proofing designed at three levels: prevention (the error can't physically happen), detection (it's caught before it moves to the next step or the customer), and alert (it's flagged so a person can act). This worksheet walks a failure mode through the three levels in order, because most fixes stop at the cheapest one — an alert — when a prevention-level fix was available for not much more.
Failure mode: the specific error being designed against, and how often it currently occurs
Current control level, if any (none / alert / detection / prevention)
Prevention option considered: can the error be made physically impossible, and at what cost
Detection option considered: how the error would be caught before the next step, and at what cost
Alert option considered: what signal would flag it, and who's responsible for acting on it
Level selected, and the reason it was chosen over a higher level (cost, feasibility — stated explicitly)
Verification: how the fix will be tested to confirm the failure mode is actually prevented, not just less frequent
// Tool 15
SMED Changeover Worksheet
Most of a changeover's cost is sequence, not equipment — and that part is free to fix.
P-02 A→B · 42:00 now · 9 steps go external · 18:40 projected
A long changeover usually isn't long because the machine is slow — it's long because half the setup work that could happen while the machine is still running gets done after it's already stopped. SMED starts by classifying every step of the current changeover as internal (only possible with the machine down) or external (can be prepped in advance), then converts as many internal steps to external as possible, before touching a single tool or spending on automation. Most of the time saved comes from that sequencing work alone — it's organizational, not capital.
Changeover being studied: from what product/setup to what, on which equipment
Full step-by-step log of the current changeover, timed, in the order it actually happens
Classification of each step: internal (machine must be stopped) or external (can be done running)
Steps identified for conversion from internal to external, and what's needed to make that happen
Steps identified for simplification: eliminating adjustment (stops/positioning), one-motion fasteners
New sequence with internal-only steps, and the projected changeover time
Actual time after implementation, measured against the baseline — not estimated
// Tool 16
Skills Matrix
If one person leaves, do you find out what breaks before or after they're gone?
5 people × 6 tasks · press setup is a singleton · C. Diaz backup due 30 Sep
Most owners can name their best technician, but few can say, in writing, who else on the team could actually cover that role tomorrow. A skills matrix lists every critical task against every person on the team and scores real proficiency — not job title — so a single point of failure shows up as a visibly empty row before it becomes a crisis. It also turns training into a plan instead of a reaction: the matrix shows exactly which gap to close next and who's ready to close it.
Critical tasks/skills listed down one axis, one row each
Team members listed across the top, one column each
Proficiency score per person per task (e.g., 0 = no exposure, 1 = trained/needs supervision, 2 = independent, 3 = can train others)
Minimum coverage target per task: how many people need to be at level 2+ for the operation to be safe
Gaps flagged where coverage falls below the target (single point of failure)
Training plan: who's being developed on which task next, and by when
Review/update date, so the matrix reflects turnover and new hires, not a snapshot from a year ago
// Tool 17
Gemba Walk Checklist
The dashboard tells you there's a problem. The floor tells you why.
Fri 15 Aug · cutoff not posted · guard near-miss closed same day · 2 escalations
A weekly report can show a number moving the wrong way without ever explaining what's actually happening at the point where the work gets done. A gemba walk is a structured visit to that exact spot — observing the process running, talking to the person doing it, and checking it against the documented standard — instead of trying to diagnose the cause from a screen. The habit tends to erode first at the top of an organization, not on the floor: the higher the level, the easier it is to trade the walk for the dashboard, which is exactly where early warning signs get missed.
Area/process to visit, and the standard it's meant to follow (what “normal” looks like)
Observation questions: is the process running as documented, and if not, where does it diverge
Questions for the person doing the work: what's slowing them down, what workaround are they using and why
Safety check: any hazard or near-miss condition observed, logged regardless of severity
Anomalies noted on the spot, with enough detail to follow up (not “check later from memory”)
Immediate actions taken during the walk vs. items escalated for follow-up
Next walk scheduled: date and whether prior action items were verified closed
// Tool 18
TOC Five Focusing Steps
Improving everything that isn't the bottleneck doesn't move your output — it just moves the queue.
Pack is the constraint · exploit first (cutoff rule) · 86→94 u/day · do not buy yet
Most operations spread improvement effort evenly across every station, which feels fair and produces almost no change in total output, because only one point in the system is actually limiting it. Theory of Constraints starts by identifying that one constraint, then works through five steps in order: exploit it without spending capital, subordinate every other step to its pace, elevate it with investment only once exploitation is exhausted, and repeat — because once a constraint breaks, the next one is already waiting somewhere else. Skipping straight to “elevate” (buy more capacity) before exploiting what's already there is the most common way to spend money on a problem that free scheduling changes would have fixed.
System/value stream being analyzed, and its current total output or throughput
Constraint identification: the step or resource limiting the whole system, with the data that proves it (not the loudest complaint)
Exploit actions: how to get maximum output from the constraint without capital spend (protect it from starvation/blocking, eliminate unnecessary setup on it)
Subordinate actions: how every other step adjusts its pace to the constraint, even if that means visible idle time elsewhere
Elevate options considered, only after exploitation is exhausted, with cost and expected throughput gain
Result: measured throughput before vs. after, on the same basis
Next constraint identified after this one is broken, so the cycle repeats instead of stopping
// Tool 19
Takt Time vs. Cycle Time
Running fast doesn't mean running at the right speed — it might just mean building inventory nobody ordered yet.
A line or a team can look busy and still be badly out of sync with what customers actually need, because “fast” and “right speed” aren't the same thing. Takt time is the pace demand requires — available time divided by customer demand for that period; cycle time is how long the process actually takes per unit. Comparing the two, station by station, shows exactly where you're the bottleneck (cycle slower than takt) and where you're overstaffed or overbuilding (cycle much faster than takt) — the same diagnosis works for a service counter or a booking calendar as it does for a production line.
Available production/service time for the period (shift length minus planned breaks/stops)
Customer demand for that same period, and its source (order history, bookings — dated)
Takt time calculation: available time ÷ demand
Cycle time measured per step/station, actual (timed), not the rated/catalog speed
Comparison chart: cycle time vs. takt time, station by station
Stations where cycle exceeds takt (bottleneck) flagged for rebalancing or capacity action
Stations where cycle is well under takt flagged for possible reallocation, not just left idle
// Talk to David
Want to see this applied to your business?
30 minutes, no pitch deck — a straight look at your operation from someone who spent twenty years running them, and what sealing the leak could be worth.