Data Strategy

Six Sigma Logistics: Defining Metrics with Global Telemetry Data Lakes

How the Define and Measure phases of Six Sigma change when every logistics system's raw records are kept in one place, and what to ask of an ingestion layer before you choose one.

Updated 11 September 2026 · first published 13 April 2026 · 8 min read

Runink Logistics Operations Team

Six Sigma Logistics: Defining Metrics with Global Telemetry Data Lakes

What are the Key Takeaways from this Executive Summary?

Quick answer

Six Sigma’s Define and Measure phases rest on a baseline everyone agrees on. In logistics that agreement is hard to get, because the warehouse system, the transport system and the yard system each record the same event in their own way. A data lake — one place that keeps every system’s raw records, unsummarized — gives the continuous improvement team a single set of numbers to work from. The harder question is what happens when two of those records contradict each other.
  • One set of records, not three: the WMS, the TMS and the YMS each hold part of the story. Keeping all of it in one place means a baseline is no longer an argument about whose export is right.
  • Measure with what the network already reports: ERPs, ELDs (electronic logging devices) and sensors already send arrivals, dwell and departures. Reading them as they arrive is what makes OTIF (on-time in-full), fill rate and dwell time measurable rather than reconstructed after the fact.
  • Specify the integration by behavior, not by logo: what matters is not which warehouse a connector targets, but what happens when two source systems disagree about the same event.

Why Do Traditional Six Sigma Implementations Fail in Modern Logistics?

Quick answer

Because the Measure phase runs on numbers pulled by hand from systems that do not agree with each other. A baseline built that way is already out of date when it lands, and nobody can say by how much.

Operations and supply chain leaders have used Lean Six Sigma for decades to find and remove defects in their networks. DMAIC — Define, Measure, Analyze, Improve, Control — is a settled method. What has changed is the number of systems a logistics defect has to be traced across, and that change lands hardest on the first two phases.

Defining a defect was simple when a supply chain was a line. A missed OTIF target, excessive yard dwell, a drayage move routed the long way round: each sat in one system. Today the same defect is spread across a carrier’s tracking feed, a customs broker’s milestones, a yard gate log and a warehouse receiving record. Small differences in how a port records an arrival turn into large differences in what final-mile performance looks like at the end of the month.

When the Measure phase depends on batch reports exported from one warehouse management system (WMS) or an ageing transportation management system (TMS), the numbers are old by the time the improvement team reads them. You cannot reduce standard deviation when the baseline itself is in doubt. A Master Black Belt cannot map a process or compute a capability index while carrier updates, customs clearance notices and warehouse receiving logs disagree about when things happened. Projects stall in Define and Measure, and the argument turns into one about the data rather than the process.


How Can Global Telemetry Data Lakes Transform the Define Phase?

Quick answer

A data lake keeps every system’s raw records in one place, so the Define phase can name a defect in terms any team can check: which lane, which stage, which measure, over which period. The precision comes from the records, not from the framework.

Define is about naming the problem, scoping the project and choosing the Critical-to-Quality (CTQ) measures the customer actually feels. In logistics those are usually fill rate, transit time and cost to serve. Naming them precisely is hard when freight moves intermodally, crosses borders under CIF or FOB terms, and changes hands three times before the final mile.

Telemetry is simply the stream of machine-generated records your equipment and your partners already produce: GPS positions, temperature readings, EDI load tenders, customs clearance milestones, gate scans. A data lake holds all of it as it arrived, not as a monthly summary. That gives Define something it rarely has — a record of what happened that was not written to support anyone’s case.

The practical effect is on the project charter. Instead of a vague objective like “reduce carrier delays,” a charter can name the lane, the season, the stage and the measure: LTL transit variance on the trans-Pacific corridor during peak season, measured at origin cross-dock dwell. This post puts no target number there, because the target belongs to your own baseline. Write the charter so that someone reading it can go to a system and check the figure. That is the test of a good problem statement, and it is easier to pass when the definition of the defect is anchored in continuous records rather than in recollection. For teams looking at specific applications, our industry use cases show where this work usually starts.


What Role Does Clean Data Ingestion Play in the Measure Phase?

Quick answer

Measure needs records that mean the same thing across sources. Ingestion is where units are reconciled, timestamps are put on one clock and missing fields are named rather than filled in. Process capability computed on top of unreconciled records is precise and wrong.

Measure establishes where you are now. For that, the numbers have to be statistically usable. If the data behind your control charts and capability figures (Cp, Cpk) is inconsistent, the rest of the DMAIC cycle inherits the error without showing it.

Raw telemetry is messy. ELDs, yard management systems (YMS) and sensors send high volumes of records in formats that differ by vendor, with gaps where a device lost signal. Ingestion is the step that checks, standardizes and validates those records before they are stored. Skip it and the data lake fills with material nobody can compute on. The variance you then observe may be coming from the process, or from the way three vendors write the word “delivered” — and you cannot tell which.

For a VP of Supply Chain, clean ingestion means a carrier timestamp from Europe lands on the same clock as a cross-dock receiving record in North America. It means detention and demurrage charges are attached to the right shipment leg rather than to the month they were invoiced in. When that work is automated, Measure stops being weeks of data gathering and becomes a reading of the network. Procurement, warehousing and transport then look at the same figures, which is usually what makes the next phase a discussion about the process.


What Does Data Lake Integration Actually Have to Solve?

Quick answer

Three things, in this order: getting logistics telemetry out of the systems that hold it, normalizing records that differ by carrier and facility, and keeping the result current enough that a Six Sigma baseline computed on Monday is still true on Friday. The third is the one that defeats most projects, because it is an operating commitment rather than a build.

Moving Six Sigma onto live data usually gets stuck on plumbing. Operations cannot wait months for IT to write custom connections to each legacy system. And when IT is spending that time maintaining fragile copy-and-reformat jobs just to produce an OTIF figure, the improvement project loses its momentum before Analyze begins.

The hard part is rarely moving the records. It is that “delivered” does not mean the same thing from two carriers, that a dwell timestamp may be taken at the gate or at the door depending on the facility, and that an EDI 214 from one forwarder carries fields another leaves blank. A pipeline that moves all of it faithfully has moved the ambiguity too, and the Six Sigma team finds out three weeks into Measure.

So the question to ask of any ingestion layer — built, bought or assembled — is not how quickly it connects. It is what it does when two sources disagree. Does the disagreement surface as a named record for someone to resolve, or does it get averaged into the baseline? A baseline that quietly absorbs contradictions will still produce a capability figure, and the figure will be wrong in a direction nobody can trace.

That is the property worth specifying before any tool is chosen: disagreements between sources must arrive as items, not as variance.


Conclusion

Quick answer

Define and Measure improve when every system’s raw records sit in one place and the definitions travel with them. An agreed baseline is what makes the later phases an argument about the process rather than about the numbers.

Lean Six Sigma has not dated. The tools used to run it have. For operations and supply chain leaders the question is no longer which method to use; it is how to get data good enough for the method to work across a global network.

Putting the records in one place closes the gaps left by separate systems, but only if the definitions travel with the data. Start by writing down, for your own network, how “on time” and “in full” are computed in each source system today, and where those definitions differ. That document is usually the real deliverable of the Measure phase. Contact the Runink team if it would help to work through it.



Sources

Six Sigma Telemetry Data Lake Runink

All articles

Runink: Data You Can Trust. Decisions You Can Defend.

Runink FACE reads the logistics records you already hold, compares them against the rules that govern them, and drafts an action for the person who owns the decision. It runs on infrastructure you control.