Pitch meetings rarely make it easy to tell data engineering vendors apart. The promises tend to sound similar: integrate your plant floor with the cloud, build high-frequency pipelines, and prepare your business for AI.
The differences become clear during delivery, when the engineers have to work with the realities of factory data rather than a clean technical brief. A generalist vendor may underestimate constraints around PLC access, struggle with legacy historians, or fall back on familiar technologies that add unnecessary cost or complexity.
Deloitte’s 2025 US Smart Manufacturing Survey found that 65–70% of manufacturers outsource roles across technology, data, and cybersecurity, while 65% rank operational risk as a first- or second-level concern in smart manufacturing initiatives. That makes it risky to bring in a partner that has to learn industrial realities on your dime.
Here are the questions, evidence, and red flags to examine before signing a contract.
Will the investment lead to measurable business outcomes?
Before architecture decisions take over the conversation, be clear about what you want to improve and how success will be measured.
That could mean reducing manual reporting, shortening production investigations, improving access to operational data, or making new sources easier to onboard. The priorities will vary, but clarity at this stage helps guide technical decisions from the start.
How will the vendor measure operational impact?
Technical metrics still matter because they show whether the solution is performing as expected. But availability, latency, and pipeline performance don't tell you whether the work is improving operations. The more important question is whether the processes are becoming faster, easier, or more reliable.
Let's look at an example. In one chemical-industry data platform project, the combination of real-time data consolidation, alerts, and predictive maintenance contributed to a 20% reduction in unplanned downtime. That’s a more useful measure of success than the number of pipelines migrated, tables created, or dashboards delivered.
For a broader look at how data supports operational improvements in this sector, see our guide to predictive analytics in manufacturing.
Evaluate a data engineering partner’s manufacturing experience
When evaluating a data engineering vendor, look for previous work involving a similar mix of systems, site count, data volumes, legacy constraints, and modernization requirements rather than focusing only on an exact subsector match.
Do they show comparable work they have put into production?
Have they taken you through one or two projects in enough detail to understand what was delivered?
There is a big difference between saying:
“We built a scalable data platform for a global manufacturer.”
and being able to point to a platform designed for hundreds of terabytes, with tens of billions of records generated from just two factories.
Ask which sources and technologies were involved, how many sites and what data volumes were supported, which parts of the existing environment had to remain in operation, and whether the result is now used in production.
Faster reporting, easier source onboarding, or shorter investigations tell you more than a description of the technology alone.
Problems during delivery are often just as revealing.
Perhaps an integration behaved differently than expected, a data-quality issue surfaced late, or a production constraint forced the team to change its approach. What changed during delivery, and why?
Who will deliver the work?
Find out who will be assigned to your project after presales. Meet the proposed architect or technical lead and, where possible, one or two senior engineers who will be directly involved in implementation.
Give them a situation from your own environment:
One site has a historian that must stay. Another runs a different MES. Both are expected to feed the same reporting platform but replacing either system is off the table. Where would you start?
You aren't asking for a finished architecture during a vendor interview. Pay attention to what the team wants to understand before proposing one. They will usually want to establish which system is the source of truth for production metrics, whether equipment identifiers and timestamps align, which data genuinely requires low latency, and whether extracting it could affect production workloads.
Before signing, confirm who remains involved after discovery and who owns the main architecture decisions during implementation.
How will they work with the systems and data you already have?
Manufacturers often modernize without disrupting systems that are already working reliably in production. The key question is when integration makes more sense and when replacement is justified.
If existing systems stay in place, the partner has to show how data from those different sources will be connected. The approach will depend on the source and the use case, whether that involves APIs, change data capture, batch ingestion, streaming, or a combination.
The vendor also has to account for how modernization will be phased. If plants are upgraded at different times, the architecture has to support old and new components running side by side without breaking reporting or downstream integrations.
How will they protect access to OT and production systems?
Operational technology, or OT, covers the hardware and software used to monitor or control physical production processes, including systems such as PLCs and SCADA. Historians and sensor infrastructure often provide the production data used for reporting and analytics.
Using that data creates an important constraint. The data platform must retrieve the required information without interfering with systems that support production. Unnecessary write permissions, poorly controlled connections, or queries that place too much load on operational systems introduce production risk.
A sound integration design should define which systems the data platform connect to, where read-only access is sufficient, how permissions and activity are controlled and logged, and how analytics workloads remain separated from production control.
How will they make data consistent and usable across plants and systems?
Bringing data from different sites into the same platform does not make it comparable automatically.
Take downtime as an example. One plant may calculate it from operator-entered reason codes, while another derives it from machine states. Both can report a KPI called “downtime” while measuring it differently.
The same problem also appears in equipment identifiers, product codes, units, and shift definitions. For that data to support cross-site reporting or analysis, the team first reconciles those differences and agrees on common definitions.

Production data requires enough context to interpret it correctly. A temperature or vibration reading means something different depending on which machine produced it, what product or batch was running, and when the reading was recorded.
How much should be standardized?
Some information belongs in a shared business standard. Common definitions for assets, products, units, and KPIs make cross-site reporting possible. Other details are specific to how an individual plant operates and do not belong in a company-wide standard.
Deciding where to draw that line requires input from both the manufacturer and the technical partner. A useful question to ask is:
How would you help us decide what to standardize centrally and what to keep site-specific?
Too little standardization makes comparisons across sites unreliable, but too much risks stripping away operational detail that local teams still rely on.
Will the data be easy to use on the plant floor?
Consistent data still has to be presented in a way that makes operational information easy to interpret. On the plant floor, operators may be monitoring several processes at once, viewing screens from a distance, or working in gloves. Important changes, warnings, and exceptions have to stand out without forcing users through dense dashboards or multiple interactions.
The visualization layer also affects technical decisions. Refresh rates, alert logic, information hierarchy, and the supporting data models all influence how useful the final interface is for operational decisions.
Do they make sound architecture decisions for your manufacturing environment?
Architecture decisions affect how well the new data platform fits the systems and constraints already in place. Poor choices risk increasing cost, making the platform harder to operate, or creating dependencies the internal team isn't prepared to manage.
How do they decide what architecture and technology to recommend?
The recommendation depends on the systems and tools already in use, the workloads the platform has to support, and any restrictions on data processing or storage locations. Operating costs, internal skills, and expected growth also affect the decision.
Those factors point to different options depending on the context:
- Microsoft Fabric when close integration with Power BI, Azure and the wider Microsoft data ecosystem is valuable.
- Databricks where Spark-centric engineering, advanced ML/AI workloads, or an existing Databricks estate make it a stronger architectural fit.
- Snowflake or AWS-based architectures when they fit the existing technology estate and workload requirements.
- On-premises or hybrid setups when production constraints limit reliance on the cloud.
For hybrid setups, our cloud consulting covers workload placement and phased modernization.
The important point is not which platform appears on the shortlist, but why it fits the manufacturer's environment and requirements.
How will the architecture hold up as requirements grow?
A design that works well today may become costly or difficult to operate as more sites, data sources, and workloads are added.
Suppose the business adds another plant, substantially more machine data, and several new analytics workloads. Ask the team to explain what that does to processing and storage, licensing and infrastructure spend, performance, onboarding effort, and ongoing maintenance.
“The platform scales” tells you very little if each additional site makes the platform disproportionately more expensive or requires significant rework.
Is the recommendation genuinely technology-neutral?
One good question is:
What would make you recommend against the platform you work with most often?
Cost, workload characteristics, hosting requirements, and internal skills are all legitimate reasons to choose something else. If none of them affects the recommendation, the technology decision deserves closer scrutiny.
What would they validate before committing to that architecture?
Before approving a larger implementation, the vendor must be able to explain which assumptions remain untested. Depending on the project, that could include:
- the volume and quality of representative data;
- any constraints on extracting data from operational systems;
- the complexity of the main integrations;
- the refresh times the proposed processing approach can support;
- the effect of expected data volumes on infrastructure or licensing costs.
Batch versus real-time processing is a good example. If supervisors review production data once per shift, second-by-second streaming adds infrastructure and expense without changing the decision. A time-sensitive alert has different requirements.
What gets left out of the first phase matters too. A team that explains why a capability doesn't belong in the initial stage is less likely to introduce extra infrastructure, licensing costs, or engineering work before the project requires it.
Can they tell when your manufacturing data isn't ready for AI?
Predictive maintenance is a good example of where AI depends heavily on the quality and context of historical data. Large volumes of sensor data may still be insufficient if maintenance events and equipment changes are poorly recorded, source data is inconsistent, or important operating context is missing. Our guide to predictive maintenance in manufacturing looks at these requirements in more detail.
An experienced vendor is willing to question whether AI is the right solution in the first place. Sometimes the data foundation requires more work. In other cases, the workflow is too inconsistent to automate or a simpler deterministic solution would solve the problem more effectively. Our article on how STX Next validates AI automation opportunities before implementation explains how we assess those questions before development begins.
Even when the architecture and use case hold up, implementation still carries risk if it disrupts reporting, integrations, or operational processes already in use.
Can they modernize your data platform without disrupting production?
Modernizing your data platform also carries migration risk. Reports, integrations, and operational processes already depend on it, so test the new setup before retiring the old one.
How will they maintain data quality, lineage, and visibility?
A reliable data platform detects when data is missing, outdated, or incorrect and makes those problems visible before they affect reporting or operational decisions.
Suppose an MES feed stops updating overnight but the reporting pipeline still completes successfully. By the morning shift meeting, the dashboard may look normal even though several hours of production data are missing. A reliable platform detects the stale source and shows which downstream reports or datasets are affected.
Data lineage becomes useful when a reported figure looks wrong. It shows how data travels from the original source through transformations and calculations into the final report or dashboard.
For example, if the reported scrap rate changes unexpectedly, engineers can trace that figure back through each step and identify where the change was introduced.
For each issue, the platform also shows which downstream reports, datasets, or processes are affected.
What happens when a critical data pipeline or source fails?
The platform also requires a defined recovery path when a source becomes unavailable, a pipeline fails midway through processing, or data arrives late. For critical pipelines, ask how failed processing is retried or replayed, how duplicate records are prevented, and what happens to data generated while a source is unavailable.
The recovery approach must also define acceptable limits for data loss and downtime and explain how service is restored after a more serious failure.
How will they prove the new data platform is ready for cutover?
For higher-risk migrations, the existing and new platforms often run in parallel before anything is retired.
A manufacturer replacing a legacy data warehouse, for example, could keep existing production reporting active while the new platform receives the same incremental data. Comparing output, scrap, downtime, or other critical measures across several reporting cycles helps identify differences before users and downstream systems switch over.
Those comparisons only work with agreed acceptance criteria. Define which outputs must match, what tolerance is acceptable, and which discrepancies would block the cutover.
The migration plan also requires a rollback path. Keep the existing platform available until the new one has passed validation and the reports, integrations, and other dependent systems have been tested.
Once the new platform is running reliably, the question becomes whether the internal team is equipped to operate and change it without depending on the vendor.
Will your team be able to own what gets built?
Your internal team doesn't have to become expert in every component, but it should have enough access, documentation, and operational knowledge to understand how the platform works and extend it.
What will your team own and understand after delivery?
Before signing, clarify four areas:
- Code and infrastructure: Your team has the access and ownership required to maintain and modify what was delivered.
- Documentation: Data flows, architecture, important decisions, and dependencies are clear enough for another engineer to follow.
- Operations: Monitoring, alerts, and runbooks are available to the people responsible after handover.
- Knowledge transfer: Internal engineers take part during delivery rather than receiving everything in a final presentation.
Knowledge transfer works better when internal engineers join architecture discussions, work alongside the delivery team on selected data flows, and begin taking responsibility before the engagement ends.
Continued support may still be useful, but as a choice rather than a technical dependency.
We cover broader models for dividing responsibilities between internal and external teams in our guide to data engineering consulting and outsourcing.
Final checks before choosing a data engineering partner
For a data engineering vendor evaluation in manufacturing, focus on the evidence behind the proposal and how well it reflects your systems, operational constraints, and business goals.

Data engineering vendor red flags to watch for
Be cautious of a partner that:
- proposes an architecture before understanding your systems and data;
- claims manufacturing expertise but has no comparable production work to show;
- proposes replacing functioning systems without explaining why;
- recommends real-time processing even when the use case doesn't require it;
- suggests essentially the same technology regardless of the environment;
- moves into full implementation without testing important assumptions;
- has no clear explanation of how the new data platform will be validated before cutover;
- leaves ownership, documentation, and knowledge transfer until the end.
Before signing, make sure the proposal reflects the systems, data constraints, and operating requirements identified during discovery.
Look for a clear architectural rationale, a record of any assumptions that still require validation, and a cutover plan that explains how the new platform will be tested before the old one is retired.
The end goal is a platform your team can operate and extend without unnecessary dependence on the vendor. If several answers are still unclear, there is important delivery risk left to resolve before signing.
If you're comparing partners now and want to pressure-test these decisions against a specific manufacturing environment, our data engineering services help assess the current environment, architecture options, and areas that need validation before a larger commitment.