Here’s a question worth asking at your next operations meeting.
If your IT systems went down right now, how long could your production floor keep operating?
Not how long your generator could run.
Not how long the machines themselves could stay powered.
How long could the plant actually produce, track, inspect, label, move, and ship good product during a manufacturing IT outage?
An hour?
Half a shift?
One day?
The answer is usually more complicated than people expect.
Because modern manufacturing has a funny contradiction.
The machines may be made of steel.
But the operation holding them together increasingly runs on data.
Your Machines Might Be Running While Your Plant Is Stopped
A CNC machine does not necessarily care whether somebody in accounting can open Outlook.
Production does care when the operator cannot retrieve the next job.
Or access the drawing.
Or confirm the revision.
Or print the traveler.
Or scan the material.
Or enter inspection data.
Or report production back into the ERP system.
A manufacturing IT outage does not have to shut every machine off simultaneously to cause serious trouble.
It can slowly remove the systems surrounding those machines until production becomes impractical—or unsafe—to continue.
Think about everything sitting between an incoming order and a truck leaving the dock.
ERP.
Scheduling.
Engineering files.
CAD/CAM.
Shop-floor terminals.
Barcode scanners.
Wi-Fi.
Quality systems.
Inventory records.
Shipping systems.
Network printers.
Label printers.
Email.
Vendor portals.
Cloud applications.
Voice systems.
File shares.
Authentication services.
If three or four of those disappear at once, somebody is going to have a bad morning.
The First 15 Minutes
At first, people usually think the problem is local.
“Try rebooting it.”
“Is Wi-Fi down?”
“Can you get into Epicor?”
“Mine isn’t working either.”
People restart computers.
Someone calls IT.
Supervisors walk around checking other areas.
Production may continue because the current jobs are already loaded.
No panic yet.
But the clock has started.
The First Hour
Now the ripple effects show up.
The next work order cannot be pulled.
A scanner cannot communicate.
Someone needs a drawing that lives on a network share.
Quality cannot enter inspection results.
A supervisor starts writing information on paper.
Shipping cannot print a label.
Purchasing cannot confirm something in the ERP system.
Meanwhile, IT is determining whether this is a failed switch, firewall problem, server failure, Internet outage, ransomware event, cloud-provider problem, or something nobody has discovered yet.
This is where a manufacturing IT outage reveals how dependent the facility has become on systems nobody thinks about when everything is working.
That is the curse of good IT.
When it works, it is invisible.
When it doesn’t, suddenly everybody knows exactly where the IT department sits.
Four Hours In
Now workarounds start breaking down.
Sure, you can write information on paper.
For a while.
But what happens when operators need new jobs?
How do you verify inventory?
Which revision of a drawing is correct?
How are completed quantities recorded?
How do you maintain traceability?
What happens to quality records?
Can shipping release product?
How do you reconcile everything entered manually once the systems return?
That last question is important.
Manual production during an outage often creates a second job after recovery: entering and verifying everything that happened while the systems were unavailable.
The plant may technically keep producing.
That does not mean there is no cost.
Could You Run an Entire Shift?
Some manufacturers can.
Many would struggle.
The difference often comes down to preparation.
A resilient manufacturer knows exactly which technology systems are critical to production.
Not just “the server.”
Which server?
Which applications?
Which network switches?
Which wireless access points?
Which databases?
Which cloud connections?
Which authentication services?
Which vendor-managed systems?
And what depends on each one?
That last part matters.
A system that looks insignificant on an IT diagram may be critical to operations.
Ask the person who once watched a $5 barcode scanner stop a $500,000 production process.
Little things can have big consequences.
The Downtime Cascade
Operations people understand cascading failures.
One station gets behind.
The next station runs out of work.
Then packaging waits.
Shipping waits.
A truck waits.
A customer waits.
Technology failures can work the same way.
Maybe the initial problem is plant Wi-Fi.
The machines themselves are fine.
But handheld scanners cannot communicate.
Material movement slows.
Inventory transactions stop.
Operators start keeping manual notes.
Supervisors improvise.
Now ERP inventory no longer matches physical inventory.
Production reporting falls behind.
The next shift walks into yesterday’s problem.
One network failure has created operational debt throughout the plant.
That is why a manufacturing IT outage is rarely measured accurately by simply asking, “How long was the server down?”
You need to ask how long the business was affected.
Build a Technology Recovery Order
Most manufacturers know which production assets matter most.
If compressor number one fails, maintenance knows what happens next.
If one particular machine goes down, production knows which jobs have to move.
Technology deserves the same planning.
Create a recovery priority list.
For example:
Tier 1: Systems required to safely operate and communicate.
Tier 2: ERP, production scheduling, authentication, networking, and critical file access.
Tier 3: Quality, warehouse, shipping, engineering, and supporting applications.
Tier 4: Less-critical administrative systems.
Your actual priorities will be different.
The important part is making those decisions before the outage.
When everything is dark at 3:30 a.m., that is a rotten time for six managers and four vendors to debate which server comes back first.
For manufacturers building or reviewing a recovery strategy, the National Institute of Standards and Technology (NIST) also provides guidance on contingency planning and recovering information systems after disruptions. Review NIST’s Contingency Planning Guide for Federal Information Systems.
That guidance is written broadly, but the basic principle applies on the shop floor too: know what is critical, know how it will be recovered, and test the plan before you need it.
Test the Backups
“We have backups.”
Good.
Have you restored them?
There is a difference between owning a fire extinguisher and knowing the thing works.
CISA recommends organizations maintain offline, encrypted backups of critical information and regularly test both backup availability and integrity in disaster recovery scenarios.
That testing should answer practical questions.
How long does an ERP recovery take?
Can you restore an entire server?
Can critical files be accessed if the primary environment is unavailable?
Are backups protected from the same ransomware event that could affect production systems?
Who knows the recovery procedure?
What happens if your primary IT provider is unavailable
You do not want the first full recovery test to happen during a real manufacturing IT outage.
Look Beyond Servers
Old-school disaster recovery focused heavily on servers.
Servers still matter.
But the modern plant has more dependencies.
Firewalls.
Switches.
Wireless access points.
Internet circuits.
Cloud platforms.
Microsoft 365.
Identity providers.
VPNs.
Industrial PCs.
Tablets.
Barcode scanners.
Printers.
Phones.
Remote vendor connections.
Sometimes losing one small piece can interrupt an entire workflow.
A manufacturer needs somebody looking at that environment as a connected production ecosystem.
That is a big part of what effective managed IT services for manufacturers should accomplish.
Not just fixing computers.
Understanding what keeps the plant moving.
Run a Tabletop Exercise
You do not have to shut the plant down to learn where the weak spots are.
Sit the right people around a table.
Operations.
IT.
Maintenance.
Engineering.
Quality.
Shipping.
Leadership.
Then give them a scenario.
“It is 6:15 Tuesday morning. The ERP system and shared drives are unavailable. IT suspects ransomware. Assume those systems will remain unavailable for the next eight hours.”
Then ask:
What stops immediately?
What can continue?
What information do operators need?
What must be available on paper?
Who contacts customers?
Who contacts vendors?
What gets disconnected?
Who decides whether production continues?
How are quality records maintained?
How do we recover manual transactions afterward?
Within 30 minutes, you will probably discover three dependencies nobody had documented.
That is the value of the exercise.
Your Recovery Plan Needs Vendors Too
Manufacturing IT gets messy because responsibility is spread around.
Your MSP handles one thing.
ERP support handles another.
Machine controls are managed by somebody else.
A cloud vendor handles something else.
During normal operations, that can be manageable.
During a major outage, it can become chaos.
Nobody wants to hear:
“That isn’t our system.”
The plant does not care whose system it is.
The line is either running or it isn’t.
Your continuity plan should identify vendor contacts, escalation procedures, responsibilities, account information, and dependencies before anybody needs them.
Ask the Question Now
You cannot eliminate every possible manufacturing IT outage.
Hardware fails.
Internet circuits get cut.
Software breaks.
People make mistakes.
Cyberattacks happen.
Storms happen.
Sometimes a $20 power supply decides Tuesday morning is its day to retire.
The goal is not pretending failure is impossible.
The goal is making sure failure does not become chaos.
Know which systems keep production moving.
Know which ones need to be restored first.
Know where the workarounds are.
Know how long those workarounds can realistically last.
Test your backups.
Document the vendors.
Practice the recovery plan.
And make sure the people responsible for your IT understand that restoring a manufacturer is not the same thing as restoring an ordinary office.
Because when a manufacturing IT outage hits, the question that matters is not whether somebody can get email.
The question is whether you can keep making product.
And if you do have to stop, how quickly can you safely start making it again?
That answer should be known before the clock starts ticking.
Want to find out where an IT outage could stop your production floor?
Request a Manufacturing IT Assessment here.







