Industrial data integration is hard for a reason that has nothing to do with integration technology. The data exists. It is being produced continuously by equipment that already works. What is missing is everything that would make a value interpretable: a tag carries a number, not an explanation of what it measures, which asset it belongs to, what shift it was recorded on, or how it relates to any other signal on the line. Every project that skips that work rediscovers it later, usually at the second site.
This guide covers the implementation sequence, in the order the decisions bind.
The first step of any Industrial DataOps platform is to enable connectivity — and the constraint is rarely the modern equipment.
What makes legacy equipment hard. Older PLCs, DCS units, loggers and historians speak proprietary protocols, expose data at inconsistent frequencies, and were never designed to be queried by anything other than their own HMI. Many sites have solved this historically by routing everything through an OPC UA server or a SCADA layer, which adds a licence, a failure point, and a translation step that discards context.
The decision that matters here. Whether connectivity is native or brokered. Native drivers read from the device; brokered access reads from something that already read from the device, which means you inherit whatever that intermediary chose to expose.
How Litmus approaches it. Litmus Edge provides native, out-of-the-box drivers for OT systems including PLCs, DCS, robotics, loggers and historians, with no dependency on purchasing OPC UA servers or SCADA.
Normalization converts disparate inputs into a consistent format. Doing it at the edge rather than after the data lands has a specific consequence: every downstream consumer receives the same shape, so you normalize once instead of once per destination.
Common mistake. Treating the cloud data platform as the place where data gets cleaned. That works for one site and one destination. At ten sites and four destinations it becomes forty transformation jobs, each of which can drift.
How companies centralize PLC and SCADA data in practice. Collect natively at each site, normalize and contextualize locally, then publish upward through governed pipelines — rather than forwarding raw tags to a central system and attempting to reconstruct meaning there. Litmus Edge Cascading provides a model for moving industrial data across sites, layers and systems without adding architectural complexity.
This is the step that determines whether anything you build survives the second site.
Industrial scale requires repeatable structure. Data models standardize machines, lines and plants so data can be reused consistently across sites: reusable models for assets, lines and facilities, with standardized attributes, relationships and hierarchies, supporting build-once, deploy-everywhere architectures.
The decision that matters here. Whether the model describes this filler or fillers. A model built around the machines in one plant will not travel, because the next plant's machines are named differently, configured differently, and possibly made by someone else.
A practical test. Before declaring the model finished, write down what would have to change to apply it at a plant you have not visited. If the answer includes tag names, the model is not finished.
Pipelines have to transform, enrich and route data reliably across edge and enterprise environments — not just move it. In practice that means defining collection, transformation, enrichment and routing as managed steps, triggered on schedules or by events, with visibility into what ran and what failed.
What to insist on. Version control. Industrial pipelines accumulate undocumented changes faster than almost any other kind of infrastructure, because they are edited under production pressure by whoever is on site. Litmus Edge Manager Git Integration brings Git-based version control, auditability and governance to distributed edge deployments — the discipline that makes a change reviewable rather than archaeological.
Governance is the step most often deferred and the one that determines whether the output is trusted.
What it covers: metadata cataloging so data assets and models are searchable and standardized, end-to-end lineage across pipelines, defined ownership and documentation, and quality validation at the point of publishing.
Why it cannot wait. An analytics output nobody can trace back to a source signal is an output nobody will act on. That is the actual reason industrial analytics projects stall at the factory floor: not that the dashboard is wrong, but that when an operator disputes it, no one can demonstrate where the number came from. A conclusion that cannot be traced cannot be defended, and a conclusion that cannot be defended does not change behaviour.
Litmus Data Catalog is purpose-built for this, and differs from enterprise data catalogs in that it operates on operational data structures rather than warehouse tables.
The same governed source has to serve different consumers in different shapes: real-time streams to edge applications, contextualized batches to cloud analytics, structured payloads to MES and enterprise systems.
Litmus integrates with the industrial stack already in place — PLCs, SCADA platforms, historians, MES, cloud infrastructure and analytics tools — and supports real-time data sharing across systems for analytics and Industrial AI.
The first deployment proves the use case. The second reveals that tags, context and structures differ at every site. Standardizing the foundation is what turns deployment two into a rollout instead of another integration project.
Reference point: a Food & Beverage manufacturer reached 95 global sites in 18 weeks using template-based rollout through Litmus Edge Manager. Niagara Bottling standardized across more than 50 plants, normalizing at the edge and streaming to Databricks.
Industrial DataOps sits alongside existing infrastructure rather than replacing it. The historian stores your data. A unified namespace organizes real-time access to it. Industrial DataOps uses both as sources while adding semantic contextualization, data quality validation, transformation logic, and governed publishing to multiple destinations.
If a vendor tells you Industrial DataOps replaces your historian, that is a migration project being sold as a data project.
Read the full explainer: What is Industrial DataOps and how it works →
The data is produced by control systems designed for control rather than analysis. Tags carry values without context, protocols vary by vendor and vintage, and the structure differs from site to site — so integration work done at one plant does not transfer to the next.
Older equipment speaks proprietary protocols, exposes data inconsistently, and was never designed to be queried externally. Routing it through an OPC UA server or SCADA layer solves access but adds cost, a failure point, and a translation step that discards context.
By collecting natively at each site, normalizing and contextualizing locally, then publishing upward through governed pipelines — rather than forwarding raw tags to a central system and reconstructing meaning there.
Most often, untraceable output. When an operator disputes a number and nobody can show where it came from, the analytics stop changing behaviour regardless of whether they were correct.
No. It uses both as sources and adds contextualization, quality validation, transformation logic and governed publishing on top.
It depends on whether you standardize before or after the second site. Published multi-site results include 95 global sites in 18 weeks using template-based rollout, which is achievable only where connectivity and context were standardized first.
