Modeling a real factory in Python: Back End logic, Front End visuals, and the HMI in between.
Almost every non-trivial program is split into two responsibilities. The Back End is the part the user never sees: it holds the data, runs the rules, and makes the decisions. In our case, the back end decides when a defective tip is rejected, when assembly fires, and when a pen is shipped. The Front End is the part the user sees and interacts with: text on a screen, buttons, charts, animations. Its only job is to present the back end's state and forward the user's intent back. A clean separation means the same back end can be reused with a web front end, a desktop GUI, or — as in industry — a panel mounted on a machine.
An HMI (Human-Machine Interface) is the specialized front end used in industrial automation. It is the touchscreen on a factory floor that lets an operator start the line, see how many pens have been shipped, watch live indicators turn green or red, and intervene if something fails. An HMI is not just decoration: it is the operator's window into the back end (often a PLC or SCADA system) and the operator's only safe channel to influence it. Three rules govern good HMI design: show only what the operator needs right now, make abnormal states visually unmistakable, and never let the screen lie about the state of the machine.
Mapping this to our project: ballpen_production_line.py is the Back End — pure logic, no pixels. ballpen_animation.py is the Front End, and because it visualizes a factory in real time with live counters and station status, it is essentially a small HMI. The two communicate through a shared data model (items, bins, counters), which is exactly how a real PLC publishes tags that an HMI subscribes to.
ballpen_production_line.py)
The back end models the factory as a set of cooperating objects. Dataclasses (Component, BallPen) describe what things are — plain data with a name, a serial number, and a quality. An Enum (Quality.OK, Quality.DEFECTIVE) replaces fragile string flags with values the interpreter can check. An Abstract Base Class (Station) defines the contract every workstation must honor: given an input, return an output or None. Concrete subclasses (ComponentMaker, QualityControl, AssemblyStation, FinalInspection, Packaging) each implement process in their own way.
The ProductionLine class is a composition, not an inheritance: it has stations, it does not act as one. Between QC and Assembly we keep collections.deque buffers — the software equivalent of a physical bin holding parts between machines. The orchestrator's run method is small because each part of the system has a single, well-defined responsibility — a textbook example of the Single Responsibility Principle.
Randomness simulates manufacturing variance: each maker has a defect rate (~5%) and assembly itself has a small failure rate (~2%). QC stations short-circuit defective items, returning None to signal "discard." This makes the line behave like a real one: the throughput is below the theoretical maximum, and the system must produce slightly more than the target to compensate for losses.
ballpen_production_line.py with random.seed(42) and confirm the report shows 5 shipped pens. Then change the seed to 1 and explain why the rejection counts differ even though the code is identical.spring, to support a retractable pen. You will need to update COMPONENT_NAMES, the BallPen dataclass, the is_complete check, and the assembly logic. Show that the report now lists makers and QC stations for the spring.defect_rate to 0.4. Predict what will happen to the number of components made per shipped pen, then run the script and compare your prediction with the actual ratio.pytest test that monkey-patches random.random to always return 0.0 (forcing every part to be defective) and asserts that the line never ships a pen — i.e., the run loop should be stopped via a maximum-attempts guard you add.Station hierarchy so that process always returns a list (possibly empty) instead of "an item or None." Explain in a comment why returning a list is more flexible if a station ever needs to emit more than one output.ballpen_animation.py)The animation is a Front End on top of the same simulation idea, written with matplotlib.animation.FuncAnimation. The fundamental pattern is the game loop: every tick we (1) clear the previous frame's transient drawings, (2) spawn new inputs, (3) advance state, (4) make decisions, and (5) render. This is the same rhythm used by every video game, every web UI library, and every industrial HMI on Earth.
Each in-flight item carries a stage field — "to_qc", "to_bin", "falling", "to_assembly" — and the renderer reads that field to decide where to draw it. This is a Finite State Machine. To synchronize the four components arriving at Assembly we use a counting semaphore (arrivals_at_assembly): when it reaches 4, we decrement by 4 and emit a finished pen. In a multithreaded version this would become an asyncio.Event or threading.Semaphore, but the logic is identical.
Crucially, the state (lists of items, bin counts, counters) is decoupled from the drawing (circles, labels, station boxes). The renderer is a pure function of the state. This is the same principle behind React, Vue, and modern HMI frameworks like Ignition or WinCC: the UI is a derivation of the data, not the source of truth. Swap matplotlib for an HTML canvas and the simulator does not change a line.
ballpen_animation.py and verify that ballpen_production_line.gif is produced. Open the GIF and identify, by visual inspection, at least one component being rejected (it should fall off the conveyor at QC).SPEED from 0.13 to 0.25 and decrease TOTAL_FRAMES to 150. Explain how throughput per frame changed and why the visual quality may suffer.rejected counter exceeds 5, change the title text to red and append " - HIGH REJECT RATE". This is your first taste of HMI alarm conditions."Bin" label with a colored bar whose height grows with the bin count. Cap it at 5 units tall, and turn the bar yellow when it reaches the top. This trains your eye to see state at a glance — the cardinal HMI design rule.SimulationState class that holds items, bins, and counters, with methods tick() and snapshot(). Have the animation only call snapshot(). Justify in a comment how this separation would let you swap the renderer for a Tkinter or web HMI without touching simulation code.A Manufacturing Execution System (MES) is the software layer that sits between the shop-floor machines (PLCs, sensors, HMIs) and the enterprise business systems (ERP, supply chain). Its job is to track, coordinate, and optimize everything that happens on the factory floor in real time: which work order is running on which machine, how many good parts have been produced, where each batch of material is, and whether the line is meeting its targets right now — not tomorrow morning when someone runs a report (MES Center).
The ISA-95 standard divides factory software into four levels. Level 1 and 2 are the physical devices and control systems (sensors, PLCs, SCADA). Level 3 is the MES — it manages operations minute-to-minute. Level 4 is the ERP (SAP, Oracle) — it manages business planning day-to-week. Without the MES in between, the ERP has no live view of the factory and the machines have no awareness of production orders.
The core functions a MES provides are:
Well-known commercial MES platforms include Siemens Opcenter (formerly SIMATIC IT), Rockwell FactoryTalk, SAP ME / SAP DMC, Tulip (no-code, popular in discrete manufacturing), and Ignition by Inductive Automation (which doubles as an HMI/SCADA platform). In semiconductors, Applied Materials PROMIS and PDF Solutions Cimetrix are industry staples. Open-source options like OpenMES and Odoo Manufacturing exist for smaller operations.
The stack you build in the next section — a Python producer writing to InfluxDB, visualized in Grafana — is a miniature MES historian. It collects the same kind of time-stamped machine data that industrial historians (OSIsoft PI, InfluxDB Enterprise, Timescale) store in production at scale. The main difference is that a real MES also enforces work orders, manages operator sign-offs, and integrates with the ERP — concerns that belong above the data-collection layer you are building.
ballpen_production_line.py, ballpen_animation.py, and the InfluxDB/Grafana stack at the correct levels and justify each placement in one sentence.run(target) call, “genealogy” maps to serial numbers on Component).Component dataclass to support this?producer.py or ballpen_production_line.py that would simulate improving it.SEMI E10 is the semiconductor-industry standard that defines how to classify every second of a machine’s time into exactly one of six mutually exclusive states (E10 Unmasked). The three states used in our demo are:
These states are the building blocks of OEE (Overall Equipment Effectiveness), the single most-watched KPI in manufacturing. OEE multiplies three ratios: Availability (how much of planned time the machine actually ran), Performance (how fast it ran versus its rated speed), and Quality (good parts out of total parts started). A world-class line targets OEE ≥ 85%. A machine stuck in Unscheduled Downtime drags Availability toward zero, collapsing OEE regardless of how fast or accurately the other stations run.
Tracking E10 states in a time-series database lets a plant engineer answer questions that a simple counter cannot: How long did the bottleneck spend in Standby yesterday afternoon? Did Unscheduled Downtime cluster around a shift change? Is Productive time trending up or down across the last week? These are exactly the questions Grafana’s state-timeline panel makes visible at a glance.
producer.py (your laptop)
| HTTP POST /api/v2/write (port 8086)
v
InfluxDB (Docker container)
| Flux queries
v
Grafana (Docker container, port 3000) --> your browser
Your script never talks to Grafana directly — it only writes to InfluxDB. Grafana pulls from InfluxDB on every refresh. That decoupling is exactly the pattern used in real factories.
Save the following as docker-compose.yml in a new folder, then run docker compose up -d in that folder:
services:
influxdb:
image: influxdb:2.7
container_name: influxdb
ports: ["8086:8086"]
environment:
DOCKER_INFLUXDB_INIT_MODE: setup
DOCKER_INFLUXDB_INIT_USERNAME: admin
DOCKER_INFLUXDB_INIT_PASSWORD: admin12345
DOCKER_INFLUXDB_INIT_ORG: lecture
DOCKER_INFLUXDB_INIT_BUCKET: factory
DOCKER_INFLUXDB_INIT_ADMIN_TOKEN: my-super-secret-token
volumes: ["influx-data:/var/lib/influxdb2"]
grafana:
image: grafana/grafana:11.0.0
container_name: grafana
ports: ["3000:3000"]
environment:
GF_SECURITY_ADMIN_PASSWORD: admin
depends_on: [influxdb]
volumes: ["grafana-data:/var/lib/grafana"]
volumes:
influx-data:
grafana-data:
InfluxDB is now on http://localhost:8086, Grafana on http://localhost:3000 (login admin / admin).
Install the client once:
pip install influxdb-client
Save the following as producer.py and run it with python producer.py. It prints one row per second and writes the same row to InfluxDB:
import time, math, random
from datetime import datetime, timezone
from influxdb_client import InfluxDBClient, Point
from influxdb_client.client.write_api import SYNCHRONOUS
URL = "http://localhost:8086"
TOKEN = "my-super-secret-token"
ORG = "lecture"
BUCKET = "factory"
client = InfluxDBClient(url=URL, token=TOKEN, org=ORG)
write = client.write_api(write_options=SYNCHRONOUS)
print("Sending data to InfluxDB. Ctrl+C to stop.")
t = 0
parts_total = 0
states = ["Productive", "Standby", "Unscheduled Downtime"]
try:
while True:
temperature = 60 + 5 * math.sin(t / 10) + random.uniform(-0.3, 0.3)
if random.random() < 0.4:
parts_total += 1
state = random.choices(states, weights=[0.85, 0.10, 0.05])[0]
point = (
Point("machine")
.tag("line", "A")
.field("temperature", temperature)
.field("parts_total", parts_total)
.field("state", state)
.time(datetime.now(timezone.utc))
)
write.write(bucket=BUCKET, record=point)
print(f"t={t:4d} T={temperature:5.2f} parts={parts_total} state={state}")
t += 1
time.sleep(1)
except KeyboardInterrupt:
print("Stopped.")
finally:
client.close()
http://localhost:3000 and log in.Flux, URL = http://influxdb:8086 (the container name, reachable inside Docker’s network), Organization = lecture, Token = my-super-secret-token, Default Bucket = factory.Left sidebar → Dashboards → New → Add visualization → pick the InfluxDB datasource.
Paste this Flux query for the temperature time-series panel:
from(bucket: "factory")
|> range(start: -5m)
|> filter(fn: (r) => r._measurement == "machine" and r._field == "temperature")
|> filter(fn: (r) => r.line == "A")
Hit Refresh — the line chart moves as long as producer.py is running. Set auto-refresh to 5s (top right) for a live feel.
For the parts produced panel: change the field filter to parts_total and pick the Stat visualization. For the E10 state panel: filter on state and pick the State timeline visualization.
producer.py, find the weights=[0.85, 0.10, 0.05] line and change it to [0.50, 0.20, 0.30]. Restart the producer and observe how the temperature panel reacts (hint: a machine in downtime typically cools down). Explain the relationship in one paragraph.state field and count rows. What percentage of time was Productive?producer.py to producer_B.py, change the tag line to .tag("line", "B"), and run both scripts simultaneously. In Grafana, duplicate the state-timeline panel and filter one on r.line == "A" and the other on r.line == "B". Describe one difference you observe between the two lines’ downtime patterns.producer.py to make Performance and Quality measurable as well.A Computerized Maintenance Management System (CMMS) is specialized software that tracks, schedules, and manages all maintenance activities for equipment and facilities. While a MES focuses on production execution, a CMMS focuses on keeping the machines running by orchestrating preventive maintenance, reactive repairs, spare parts inventory, and maintenance history. It is the bridge between operations and reliability engineering.
Unplanned downtime is the enemy of OEE. A single unexpected machine failure can halt an entire production line and cost thousands per hour. A CMMS minimizes this by shifting maintenance from reactive ("the machine broke, fix it now!") to preventive ("the machine is due for maintenance Tuesday night during the scheduled stop") or even predictive ("vibration sensors show early bearing wear; schedule maintenance before failure"). This shift improves equipment availability, extends asset life, and reduces emergency repair costs.
In the ball-pen example, SEMI E10 Scheduled Downtime is maintenance initiated by a CMMS work order. When the system knows in advance that maintenance is coming, it can pre-position materials, notify the scheduler, and avoid unexpected failures that would appear as Unscheduled Downtime — the enemy of OEE.
Commercial platforms include Maximo (IBM, enterprise-scale), SAP Plant Maintenance (PM), Infor EAM, Aspen Mtx, and Computerized Relative Allocation of Facilities (CRAFT, used in hospitals and facilities). Mid-market and discrete manufacturers often use UpKeep, Fiix (now part of IBM), or Dude Solutions. Open-source options like MINTO and OpenMAINT support smaller operations.
For more details, see OSApiens CMMS, which offers a comprehensive overview of CMMS platforms, best practices, and ROI calculations.
parts_total run hours) to the CMMS (e.g., to trigger a PM task based on cycle count) and why they cannot live in one database.producer.py, add a field to record whether the line is in Scheduled Downtime (CMMS-initiated maintenance). Propose how you would trigger a scheduled downtime window and how the OEE calculation should treat that state versus Unscheduled Downtime.In the ball-pen simulation, quality control is rule-based: each component has a Quality flag (OK or DEFECTIVE), set randomly. In a real factory, quality is measured by inspection. Historically, this meant a human operator with a magnifying glass checking for scratches, misalignment, or dimensional errors — slow, inconsistent, and prone to fatigue. Computer vision and artificial intelligence have transformed QC into a high-speed, consistent, and measurable process that runs 24/7 without tiring.
Computer vision is the ability of a machine to interpret images and video. A camera on the production line captures a photo of each component or assembly. Machine learning algorithms (often deep neural networks trained on thousands of examples) instantly classify the image as pass or fail, detect specific defects (cracks, discoloration, missing features), and measure dimensions with micron precision. Unlike traditional rule-based thresholds, AI models learn subtle patterns that no engineer could codify by hand.
In our ball-pen line, the QualityControl station randomly rejects ~5% of inputs. With vision-based AI QC:
AI QC is not magic. Challenges include: lighting and image quality must be consistent (bad photos train bad models), the model must be retrained when product design changes, rare defects are hard to learn (you need many examples), and false positives can trigger unnecessary rejections that waste parts. The solution is human-in-the-loop: controversial images are flagged for human review, with the decision fed back into model retraining. This hybrid approach combines AI speed with human judgment.
Electronics manufacturing uses AI to inspect solder joints, component placement, and circuit board traces. Automotive plants inspect painted surfaces, weld integrity, and final assembly fit. Food & beverage uses vision to sort defective products, detect foreign objects, and verify label placement. Pharmaceuticals employ vision to count pills, check seal integrity, and detect vial cracks.
For deeper insight into computer vision in manufacturing, see Tupl: Computer Vision in Manufacturing, which covers real implementations, ROI, and case studies.
QualityControl.process() method, replace the random rejection logic with a deterministic rule: "reject if the component is from a batch older than 30 days." How would you track batch age in the Component dataclass to support this? What historical data from an AI vision system would make this rule possible?