An industrial DataOps reference architecture is a structured framework that defines where and how industrial data should be collected, processed, stored, and analyzed across four distinct tiers: edge, site, enterprise, and cloud. The answer to whether industrial data should be processed at the edge or in the cloud is not binary—it depends on latency requirements, bandwidth constraints, governance needs, and the specific workload. What you need for Industrial DataOps is a tiered architecture that places the right processing at the right layer, enabling real-time machine control at the edge while reserving advanced analytics and AI workloads for cloud environments with greater computational resources.
This reference architecture serves as the blueprint for manufacturers and industrial organizations seeking to standardize data flow from shop-floor devices to enterprise systems. Rather than treating edge and cloud as competing options, a well-designed industrial DataOps reference architecture treats them as complementary layers, each optimized for specific data processing requirements. The sections that follow provide tier-by-tier guidance to help solutions architects, transformation leaders, and integration engineers make informed placement decisions that balance performance, cost, and scalability.
An industrial DataOps reference architecture is a standardized blueprint that maps how operational technology (OT) data flows from machines and sensors through successive processing layers—edge, site, enterprise, and cloud—to enable analytics, automation, and AI-driven decision-making. It provides a common framework for organizations to design scalable, repeatable data infrastructure across multiple facilities.
The architecture addresses a fundamental challenge in industrial environments: data originates in highly distributed, often disconnected systems—PLCs, SCADA systems, historians, and sensors—that were never designed to share information with modern analytics platforms. A dataops architecture bridges this gap by defining clear boundaries for where data is collected, contextualized, aggregated, and analyzed.
Unlike a Unified Namespace (UNS), which provides a single semantic layer for organizing and accessing industrial data, an industrial DataOps reference architecture encompasses the complete data lifecycle. The UNS is a critical component within this architecture, but the broader framework also addresses data pipeline architecture, governance, storage tiering, and workload placement. For a deeper exploration of how UNS fits within this framework, see the Industrial DataOps and UNS reference architecture.
The four-tier model—edge, site, enterprise, and cloud—reflects the physical and logical realities of industrial operations. Each tier has distinct characteristics in terms of latency tolerance, connectivity reliability, computational capacity, and data volume. A modern data architecture for industrial environments must account for all four layers to avoid bottlenecks, reduce unnecessary data transfer costs, and ensure that time-sensitive decisions happen where they need to happen.
The edge tier is where industrial data originates and where the most time-sensitive processing must occur. This layer encompasses the devices, gateways, and embedded compute resources located directly at or near industrial equipment—on the machine, beside the production line, or within the control cabinet.
Data processing at the edge tier serves three primary purposes: real-time control and automation, data filtering and reduction, and initial contextualization. When a sensor detects an anomaly that requires immediate action—such as a temperature spike that could damage equipment—the response must happen in milliseconds. Sending that data to the cloud and waiting for a response is not viable when latency requirements are measured in single-digit milliseconds.
Edge-tier processing also addresses bandwidth constraints. Industrial equipment can generate massive volumes of raw data—vibration sensors alone may produce thousands of readings per second. Transmitting all of this data to higher tiers would overwhelm network capacity and incur significant costs. Instead, edge processing filters, aggregates, and compresses data before forwarding only what is necessary for upstream analysis.
The dataops tools deployed at this tier typically include edge gateways, protocol converters, and lightweight analytics engines capable of running on constrained hardware. These tools must support the diverse protocols found in industrial environments—OPC UA, Modbus, MQTT, EtherNet/IP, and proprietary formats—while normalizing data into consistent structures. For a detailed look at how edge analytics operates at the machine level, explore edge analytics with Litmus Edge.
AI agents at the edge tier focus on inference rather than training. Pre-trained models for anomaly detection, predictive maintenance, or quality inspection can run locally, enabling intelligent automation without cloud dependency. This is particularly critical in environments where connectivity is intermittent or where security policies prohibit external data transmission.
The site tier represents the plant or facility level, where data from multiple edge nodes is aggregated, correlated, and analyzed within the context of a single production environment. This layer bridges the gap between machine-level operations and enterprise-wide visibility.
At this tier, data from dozens or hundreds of edge devices converges into a unified view of plant operations. The site tier is where cross-machine analytics become possible—correlating data from upstream and downstream equipment to identify bottlenecks, optimize throughput, or detect quality issues that span multiple process steps.
A Unified Namespace architecture is particularly valuable at the site tier. By organizing all plant data into a consistent, hierarchical structure, the UNS enables applications and users to access information without needing to understand the underlying source systems. This abstraction layer simplifies integration and accelerates the deployment of new analytics use cases. Learn more about embracing the Unified Namespace architecture with Litmus Edge.
Site-tier infrastructure typically includes on-premises servers, local historians, and plant-level data platforms. These systems provide the computational capacity for more sophisticated analytics than edge devices can support, while maintaining the low-latency access required for operational decision-making. Batch processing, shift-level reporting, and local dashboards are common workloads at this layer.
The site tier also serves as a buffer for cloud connectivity. When network connections to enterprise or cloud systems are unavailable—whether due to planned maintenance, connectivity issues, or security policies—the site tier continues to operate independently. Data is stored locally and synchronized when connectivity is restored, ensuring continuity of operations regardless of external dependencies.
The enterprise tier aggregates data from multiple sites to enable organization-wide visibility, standardization, and governance. This layer is where industrial management data systems converge with IT infrastructure to support strategic decision-making.
At this tier, data from individual plants is harmonized into consistent formats, enabling meaningful comparisons across facilities. A manufacturer with ten plants can analyze overall equipment effectiveness (OEE) across all locations, identify best practices at high-performing sites, and replicate successful configurations elsewhere. Without enterprise-tier integration, each plant operates as an isolated data silo.
Data governance becomes critical at the enterprise tier. This includes defining data ownership, establishing quality standards, managing access controls, and ensuring compliance with regulatory requirements. The enterprise tier is where master data management intersects with operational data, linking production information to business systems like ERP and supply chain platforms.
The data platform architecture at this tier must support diverse integration patterns. Some data flows in real-time streams for operational dashboards, while other data moves in scheduled batches for reporting and compliance. The enterprise tier must accommodate both patterns while maintaining data lineage and auditability. For organizations evaluating connectivity options across systems, the Litmus integrations library provides a comprehensive view of supported protocols and platforms.
Enterprise-tier infrastructure may be deployed on-premises in corporate data centers, in private clouds, or in hybrid configurations. The choice depends on factors including data residency requirements, existing IT investments, and the organization's cloud strategy. Regardless of deployment model, the enterprise tier serves as the authoritative source for cross-facility analytics and the gateway to cloud-based advanced analytics.
The cloud tier provides virtually unlimited storage capacity, elastic compute resources, and access to advanced analytics and AI services that would be impractical to deploy at lower tiers. This layer is where industrial data realizes its full potential for strategic insight and innovation.
Cloud infrastructure excels at workloads that require massive computational resources, long-term data retention, or access to specialized services. Training machine learning models on years of historical production data, running complex simulations, or applying advanced computer vision algorithms—these tasks benefit from the scale and flexibility that cloud platforms provide.
A data lakehouse architecture is increasingly common at the cloud tier, combining the flexibility of data lakes with the performance and governance capabilities of data warehouses. This approach allows organizations to store raw operational data alongside curated, analytics-ready datasets, supporting both exploratory analysis and production reporting from a single platform.
The cloud tier is also where industrial data integrates with broader enterprise analytics ecosystems. Production data can be combined with supply chain information, customer demand signals, and financial data to enable end-to-end business optimization. For organizations building toward industrial AI, the cloud tier provides the foundation for building a modern industrial data foundation.
Integration between edge infrastructure and cloud platforms is essential for realizing this value. Partnerships between industrial DataOps platforms and cloud providers streamline data ingestion, transformation, and analysis. As an example, the integration between Litmus and Databricks demonstrates how edge-collected data can flow seamlessly into cloud-based industrial AI workflows.
AI workloads at the cloud tier focus on model training, experimentation, and advanced inference that exceeds edge capabilities. Once models are trained and validated, they can be deployed back to edge or site tiers for real-time inference, creating a continuous improvement loop that leverages the strengths of each architectural layer.
Placement decisions should be driven by four primary factors: latency requirements, bandwidth and cost constraints, data governance needs, and computational complexity. Each factor points toward different tiers for different workloads.
Latency requirements are the most decisive factor. If a process requires sub-second response times—machine control, safety systems, or real-time quality inspection—processing must occur at the edge. As latency tolerance increases to seconds or minutes, site-tier processing becomes viable. Enterprise and cloud tiers are appropriate when latency tolerance extends to hours or longer.
Bandwidth and cost constraints favor processing data as close to the source as possible. Transmitting raw, high-frequency sensor data to the cloud is expensive and often unnecessary. Edge and site tiers should filter, aggregate, and compress data before transmission, sending only the information required for upstream analysis. This reduces network costs and avoids overwhelming higher-tier systems with data that provides no additional value.
Data governance needs may require certain data to remain on-premises or within specific geographic boundaries. Regulatory requirements, customer contracts, or internal security policies may prohibit sending production data to public cloud environments. In these cases, enterprise-tier or site-tier processing provides the necessary control while still enabling analytics and optimization.
Computational complexity determines whether a workload can run on constrained edge hardware or requires the resources available at higher tiers. Simple threshold-based alerts and basic aggregations run efficiently at the edge. Machine learning inference with pre-trained models can often run at edge or site tiers. Model training, complex simulations, and large-scale historical analysis require cloud-tier resources.
The optimal architecture is rarely all-edge or all-cloud. Instead, it distributes workloads across tiers based on these factors, ensuring that each processing task occurs at the layer best suited to its requirements. For a deeper exploration of balancing edge and cloud placement, see finding the best of both worlds.
Workload Type | Recommended Tier | Key Driver |
Machine control and safety | Edge | Sub-millisecond latency |
Anomaly detection (inference) | Edge or Site | Low latency, local action |
Shift-level reporting | Site | Plant-level aggregation |
Cross-facility benchmarking | Enterprise | Multi-site data integration |
ML model training | Cloud | Computational scale |
Long-term data archival | Cloud | Storage cost efficiency |
Air-gapped and offline environments present additional constraints. In facilities without external connectivity—common in defense, critical infrastructure, and highly regulated industries—all processing must occur at edge, site, or enterprise tiers. The architecture must be designed for complete autonomy, with no dependency on cloud services for operational functionality.
Designing and implementing an industrial DataOps reference architecture requires a dataops platform that spans all four tiers while providing the flexibility to adapt to diverse industrial environments. The architecture must connect legacy OT systems, support modern protocols, and integrate with enterprise and cloud platforms without requiring wholesale replacement of existing infrastructure.
Litmus provides a comprehensive platform for industrial DataOps, enabling organizations to collect, contextualize, and deliver operational data from edge to cloud. The platform supports connectivity to hundreds of industrial protocols, ensuring that data from PLCs, SCADA systems, historians, and sensors can be unified regardless of vendor or vintage.
At the edge tier, Litmus Edge provides real-time data collection, protocol conversion, and local analytics on industrial-grade hardware. At site and enterprise tiers, the platform aggregates data into a Unified Namespace, enabling consistent access across applications and users. Cloud integrations deliver contextualized data to platforms including Databricks, AWS, Azure, and Google Cloud for advanced analytics and AI workloads.
For solutions architects designing scalable infrastructure, transformation leaders evaluating platform investments, and integration engineers implementing connectivity, the Litmus Architecture Hub provides reference architectures, deployment patterns, and technical guidance to accelerate your industrial DataOps journey.
An industrial DataOps reference architecture is a standardized framework that defines how operational data flows from industrial equipment through edge, site, enterprise, and cloud tiers for processing, storage, and analysis. It provides a blueprint for designing scalable data infrastructure that connects OT systems with modern analytics platforms. The architecture addresses data collection, contextualization, governance, and workload placement across all layers.
The four key layers are edge, site, enterprise, and cloud. The edge tier handles real-time processing at the machine level. The site tier aggregates data at the plant level for local analytics. The enterprise tier integrates data across multiple facilities and enforces governance. The cloud tier provides scalable storage, advanced analytics, and AI model training capabilities.
The optimal placement depends on latency requirements, bandwidth constraints, governance needs, and computational complexity. Time-sensitive workloads like machine control must process at the edge. Plant-level analytics belong at the site tier. Cross-facility analysis occurs at the enterprise tier. Advanced AI training and long-term storage are best suited for the cloud. Most organizations require a hybrid approach that distributes workloads across all tiers.
Industrial DataOps encompasses the complete data lifecycle—collection, contextualization, processing, storage, and analysis—across all architectural tiers. A Unified Namespace is a specific component within this architecture that provides a single semantic layer for organizing and accessing industrial data. The UNS simplifies data access but does not address the full scope of data pipeline architecture, governance, or workload placement that industrial DataOps covers.
An industrial DataOps architecture connects PLCs, SCADA systems, historians, sensors, MES platforms, ERP systems, and other OT and IT data sources. It supports diverse industrial protocols including OPC UA, Modbus, MQTT, EtherNet/IP, and proprietary formats. The architecture normalizes data from these disparate sources into consistent structures for unified analysis.
Industrial DataOps bridges the gap between operational technology and information technology by providing a common data infrastructure that both domains can access. It contextualizes raw OT data into formats that IT systems and business applications can consume. This convergence enables use cases like integrating production data with ERP systems, combining operational metrics with financial analysis, and applying enterprise AI capabilities to shop-floor optimization.
The primary challenges include connecting legacy equipment that uses proprietary or outdated protocols, managing the volume and velocity of industrial data without overwhelming network and storage resources, ensuring data quality and consistency across diverse sources, and balancing security requirements with the need for data accessibility. Organizations must also address organizational challenges, including aligning OT and IT teams around shared data standards and governance practices.
