Production Line

Modeling a real factory in Python: Back End logic, Front End visuals, and the HMI in between.

1. Software Architecture: Front End, Back End, and HMI

Theory (Deeper Concepts)

Siemens HMI

Almost every non-trivial program is split into two responsibilities. The Back End is the part the user never sees: it holds the data, runs the rules, and makes the decisions. In our case, the back end decides when a defective tip is rejected, when assembly fires, and when a pen is shipped. The Front End is the part the user sees and interacts with: text on a screen, buttons, charts, animations. Its only job is to present the back end's state and forward the user's intent back. A clean separation means the same back end can be reused with a web front end, a desktop GUI, or — as in industry — a panel mounted on a machine.

An HMI (Human-Machine Interface) is the specialized front end used in industrial automation. It is the touchscreen on a factory floor that lets an operator start the line, see how many pens have been shipped, watch live indicators turn green or red, and intervene if something fails. An HMI is not just decoration: it is the operator's window into the back end (often a PLC or SCADA system) and the operator's only safe channel to influence it. Three rules govern good HMI design: show only what the operator needs right now, make abnormal states visually unmistakable, and never let the screen lie about the state of the machine.

Mapping this to our project: ballpen_production_line.py is the Back End — pure logic, no pixels. ballpen_animation.py is the Front End, and because it visualizes a factory in real time with live counters and station status, it is essentially a small HMI. The two communicate through a shared data model (items, bins, counters), which is exactly how a real PLC publishes tags that an HMI subscribes to.

Exercises (Deeper Focus)

  1. List three responsibilities of a Back End and three of a Front End in any system you use daily (e.g., your bank app, an online store). Identify which side handles authentication, which renders the cart, and which validates a price.
  2. Find a photo of a real industrial HMI online and identify at least four pieces of information it shows the operator. For each, classify it as status, alarm, metric, or control.
  3. Explain in one paragraph why coupling the Back End and Front End tightly (e.g., letting the UI directly read a sensor) is dangerous in a factory. Use the words safety, latency, and testability.

2. Ball Pen Production - The Back End (ballpen_production_line.py)

Theory (Deeper Concepts)

BIC Cristal Pen Schema

The back end models the factory as a set of cooperating objects. Dataclasses (Component, BallPen) describe what things are — plain data with a name, a serial number, and a quality. An Enum (Quality.OK, Quality.DEFECTIVE) replaces fragile string flags with values the interpreter can check. An Abstract Base Class (Station) defines the contract every workstation must honor: given an input, return an output or None. Concrete subclasses (ComponentMaker, QualityControl, AssemblyStation, FinalInspection, Packaging) each implement process in their own way.

The ProductionLine class is a composition, not an inheritance: it has stations, it does not act as one. Between QC and Assembly we keep collections.deque buffers — the software equivalent of a physical bin holding parts between machines. The orchestrator's run method is small because each part of the system has a single, well-defined responsibility — a textbook example of the Single Responsibility Principle.

Randomness simulates manufacturing variance: each maker has a defect rate (~5%) and assembly itself has a small failure rate (~2%). QC stations short-circuit defective items, returning None to signal "discard." This makes the line behave like a real one: the throughput is below the theoretical maximum, and the system must produce slightly more than the target to compensate for losses.

Exercises (Deeper Focus)

  1. Run ballpen_production_line.py with random.seed(42) and confirm the report shows 5 shipped pens. Then change the seed to 1 and explain why the rejection counts differ even though the code is identical.
  2. Add a fifth component, spring, to support a retractable pen. You will need to update COMPONENT_NAMES, the BallPen dataclass, the is_complete check, and the assembly logic. Show that the report now lists makers and QC stations for the spring.
  3. Raise every component's defect_rate to 0.4. Predict what will happen to the number of components made per shipped pen, then run the script and compare your prediction with the actual ratio.
  4. Write a pytest test that monkey-patches random.random to always return 0.0 (forcing every part to be defective) and asserts that the line never ships a pen — i.e., the run loop should be stopped via a maximum-attempts guard you add.
  5. Refactor the Station hierarchy so that process always returns a list (possibly empty) instead of "an item or None." Explain in a comment why returning a list is more flexible if a station ever needs to emit more than one output.

3. Ball Pen Animation - The HMI (ballpen_animation.py)

Theory (Deeper Concepts)

The animation is a Front End on top of the same simulation idea, written with matplotlib.animation.FuncAnimation. The fundamental pattern is the game loop: every tick we (1) clear the previous frame's transient drawings, (2) spawn new inputs, (3) advance state, (4) make decisions, and (5) render. This is the same rhythm used by every video game, every web UI library, and every industrial HMI on Earth.

Each in-flight item carries a stage field — "to_qc", "to_bin", "falling", "to_assembly" — and the renderer reads that field to decide where to draw it. This is a Finite State Machine. To synchronize the four components arriving at Assembly we use a counting semaphore (arrivals_at_assembly): when it reaches 4, we decrement by 4 and emit a finished pen. In a multithreaded version this would become an asyncio.Event or threading.Semaphore, but the logic is identical.

Crucially, the state (lists of items, bin counts, counters) is decoupled from the drawing (circles, labels, station boxes). The renderer is a pure function of the state. This is the same principle behind React, Vue, and modern HMI frameworks like Ignition or WinCC: the UI is a derivation of the data, not the source of truth. Swap matplotlib for an HTML canvas and the simulator does not change a line.

Exercises (Deeper Focus)

  1. Run ballpen_animation.py and verify that ballpen_production_line.gif is produced. Open the GIF and identify, by visual inspection, at least one component being rejected (it should fall off the conveyor at QC).
  2. Increase SPEED from 0.13 to 0.25 and decrease TOTAL_FRAMES to 150. Explain how throughput per frame changed and why the visual quality may suffer.
  3. Add an alarm behavior worthy of an HMI: when the rejected counter exceeds 5, change the title text to red and append " - HIGH REJECT RATE". This is your first taste of HMI alarm conditions.
  4. Replace the static "Bin" label with a colored bar whose height grows with the bin count. Cap it at 5 units tall, and turn the bar yellow when it reaches the top. This trains your eye to see state at a glance — the cardinal HMI design rule.
  5. Decouple the simulation from the renderer cleanly: extract a SimulationState class that holds items, bins, and counters, with methods tick() and snapshot(). Have the animation only call snapshot(). Justify in a comment how this separation would let you swap the renderer for a Tkinter or web HMI without touching simulation code.

4. Manufacturing Execution Systems (MES)

Theory (Deeper Concepts)

A Manufacturing Execution System (MES) is the software layer that sits between the shop-floor machines (PLCs, sensors, HMIs) and the enterprise business systems (ERP, supply chain). Its job is to track, coordinate, and optimize everything that happens on the factory floor in real time: which work order is running on which machine, how many good parts have been produced, where each batch of material is, and whether the line is meeting its targets right now — not tomorrow morning when someone runs a report (MES Center).

The ISA-95 standard divides factory software into four levels. Level 1 and 2 are the physical devices and control systems (sensors, PLCs, SCADA). Level 3 is the MES — it manages operations minute-to-minute. Level 4 is the ERP (SAP, Oracle) — it manages business planning day-to-week. Without the MES in between, the ERP has no live view of the factory and the machines have no awareness of production orders.

The core functions a MES provides are:

  • Work order management — dispatches production orders to the right machine at the right time and tracks their progress to completion.
  • Genealogy & traceability — records which materials, operators, and machines touched every unit, so a recall can be scoped to exactly the affected batch.
  • Quality management — captures inline inspection results, triggers holds on suspect lots, and feeds SPC (Statistical Process Control) charts.
  • Performance monitoring (OEE) — aggregates E10 machine states, cycle times, and yield data into OEE dashboards visible to operators and managers simultaneously.
  • Data collection & historians — stores time-stamped process data (temperatures, pressures, speeds) that engineers mine for root-cause analysis and process improvement.

Well-known commercial MES platforms include Siemens Opcenter (formerly SIMATIC IT), Rockwell FactoryTalk, SAP ME / SAP DMC, Tulip (no-code, popular in discrete manufacturing), and Ignition by Inductive Automation (which doubles as an HMI/SCADA platform). In semiconductors, Applied Materials PROMIS and PDF Solutions Cimetrix are industry staples. Open-source options like OpenMES and Odoo Manufacturing exist for smaller operations.

The stack you build in the next section — a Python producer writing to InfluxDB, visualized in Grafana — is a miniature MES historian. It collects the same kind of time-stamped machine data that industrial historians (OSIsoft PI, InfluxDB Enterprise, Timescale) store in production at scale. The main difference is that a real MES also enforces work orders, manages operator sign-offs, and integrates with the ERP — concerns that belong above the data-collection layer you are building.

Exercises (Deeper Focus)

  1. Look up the ISA-95 standard (a one-page summary is enough). Draw a four-level pyramid showing which software category lives at each level. Place our ballpen_production_line.py, ballpen_animation.py, and the InfluxDB/Grafana stack at the correct levels and justify each placement in one sentence.
  2. Choose one commercial MES from the list above and find its product page. Identify three features it advertises. For each feature, name the equivalent concept in our ball-pen simulation (e.g., “work order” maps to the run(target) call, “genealogy” maps to serial numbers on Component).
  3. A recall is issued for pens assembled between 14:00 and 15:00 yesterday using barrel components from lot “LOT-42”. Describe in plain English the data a MES would need to store — and what queries it would run — to produce a list of affected serial numbers. What fields would you add to the Component dataclass to support this?
  4. OEE = Availability × Performance × Quality. A line runs 8 hours, has 45 minutes of unscheduled downtime, produces at 90% of rated speed, and yields 96% good parts. Calculate the OEE. Then identify which factor is the weakest link and suggest one change to producer.py or ballpen_production_line.py that would simulate improving it.

5. Live Factory Telemetry: Python → InfluxDB → Grafana

Theory (Deeper Concepts)

E10 Machine State — What It Is and Why It Matters

SEMI E10 is the semiconductor-industry standard that defines how to classify every second of a machine’s time into exactly one of six mutually exclusive states (E10 Unmasked). The three states used in our demo are:

  • Productive — the machine is processing parts. This is the only state that generates revenue.
  • Standby — the machine is ready but waiting for material, an operator, or a downstream station. The line is running; this machine is idle.
  • Unscheduled Downtime — the machine failed unexpectedly. No parts are moving and no recovery was planned.

These states are the building blocks of OEE (Overall Equipment Effectiveness), the single most-watched KPI in manufacturing. OEE multiplies three ratios: Availability (how much of planned time the machine actually ran), Performance (how fast it ran versus its rated speed), and Quality (good parts out of total parts started). A world-class line targets OEE ≥ 85%. A machine stuck in Unscheduled Downtime drags Availability toward zero, collapsing OEE regardless of how fast or accurately the other stations run.

Tracking E10 states in a time-series database lets a plant engineer answer questions that a simple counter cannot: How long did the bottleneck spend in Standby yesterday afternoon? Did Unscheduled Downtime cluster around a shift change? Is Productive time trending up or down across the last week? These are exactly the questions Grafana’s state-timeline panel makes visible at a glance.

How the Pieces Connect

producer.py (your laptop)
      |  HTTP POST /api/v2/write   (port 8086)
      v
   InfluxDB  (Docker container)
      |  Flux queries
      v
   Grafana   (Docker container, port 3000) --> your browser

Your script never talks to Grafana directly — it only writes to InfluxDB. Grafana pulls from InfluxDB on every refresh. That decoupling is exactly the pattern used in real factories.

Step 1 — Start InfluxDB and Grafana

Save the following as docker-compose.yml in a new folder, then run docker compose up -d in that folder:

services:
  influxdb:
    image: influxdb:2.7
    container_name: influxdb
    ports: ["8086:8086"]
    environment:
      DOCKER_INFLUXDB_INIT_MODE: setup
      DOCKER_INFLUXDB_INIT_USERNAME: admin
      DOCKER_INFLUXDB_INIT_PASSWORD: admin12345
      DOCKER_INFLUXDB_INIT_ORG: lecture
      DOCKER_INFLUXDB_INIT_BUCKET: factory
      DOCKER_INFLUXDB_INIT_ADMIN_TOKEN: my-super-secret-token
    volumes: ["influx-data:/var/lib/influxdb2"]

  grafana:
    image: grafana/grafana:11.0.0
    container_name: grafana
    ports: ["3000:3000"]
    environment:
      GF_SECURITY_ADMIN_PASSWORD: admin
    depends_on: [influxdb]
    volumes: ["grafana-data:/var/lib/grafana"]

volumes:
  influx-data:
  grafana-data:

InfluxDB is now on http://localhost:8086, Grafana on http://localhost:3000 (login admin / admin).

Step 2 — Python Script That Generates Data

Install the client once:

pip install influxdb-client

Save the following as producer.py and run it with python producer.py. It prints one row per second and writes the same row to InfluxDB:

import time, math, random
from datetime import datetime, timezone
from influxdb_client import InfluxDBClient, Point
from influxdb_client.client.write_api import SYNCHRONOUS

URL    = "http://localhost:8086"
TOKEN  = "my-super-secret-token"
ORG    = "lecture"
BUCKET = "factory"

client = InfluxDBClient(url=URL, token=TOKEN, org=ORG)
write  = client.write_api(write_options=SYNCHRONOUS)

print("Sending data to InfluxDB. Ctrl+C to stop.")
t = 0
parts_total = 0
states = ["Productive", "Standby", "Unscheduled Downtime"]

try:
    while True:
        temperature = 60 + 5 * math.sin(t / 10) + random.uniform(-0.3, 0.3)
        if random.random() < 0.4:
            parts_total += 1
        state = random.choices(states, weights=[0.85, 0.10, 0.05])[0]

        point = (
            Point("machine")
            .tag("line", "A")
            .field("temperature", temperature)
            .field("parts_total", parts_total)
            .field("state", state)
            .time(datetime.now(timezone.utc))
        )
        write.write(bucket=BUCKET, record=point)
        print(f"t={t:4d}  T={temperature:5.2f}  parts={parts_total}  state={state}")
        t += 1
        time.sleep(1)
except KeyboardInterrupt:
    print("Stopped.")
finally:
    client.close()

Step 3 — Wire Grafana to InfluxDB (One-Time)

  1. Open http://localhost:3000 and log in.
  2. Left sidebar → Connections → Data sources → Add data source → InfluxDB.
  3. Set: Query language = Flux, URL = http://influxdb:8086 (the container name, reachable inside Docker’s network), Organization = lecture, Token = my-super-secret-token, Default Bucket = factory.
  4. Click Save & test — you should see “datasource is working”.

Step 4 — Build a Panel

Left sidebar → Dashboards → New → Add visualization → pick the InfluxDB datasource.

Paste this Flux query for the temperature time-series panel:

from(bucket: "factory")
  |> range(start: -5m)
  |> filter(fn: (r) => r._measurement == "machine" and r._field == "temperature")
  |> filter(fn: (r) => r.line == "A")

Hit Refresh — the line chart moves as long as producer.py is running. Set auto-refresh to 5s (top right) for a live feel.

For the parts produced panel: change the field filter to parts_total and pick the Stat visualization. For the E10 state panel: filter on state and pick the State timeline visualization.

Exercises (Deeper Focus)

  1. Complete Steps 1–4 above from scratch. Take a screenshot of all three panels running live. On the state timeline, identify at least one Unscheduled Downtime interval and estimate how many seconds it lasted.
  2. In producer.py, find the weights=[0.85, 0.10, 0.05] line and change it to [0.50, 0.20, 0.30]. Restart the producer and observe how the temperature panel reacts (hint: a machine in downtime typically cools down). Explain the relationship in one paragraph.
  3. Add a fourth Grafana panel showing a pie chart of total time spent in each E10 state over the last 15 minutes. Write the Flux query needed to group by the state field and count rows. What percentage of time was Productive?
  4. Add a second production line: copy producer.py to producer_B.py, change the tag line to .tag("line", "B"), and run both scripts simultaneously. In Grafana, duplicate the state-timeline panel and filter one on r.line == "A" and the other on r.line == "B". Describe one difference you observe between the two lines’ downtime patterns.
  5. Calculate the simulated OEE for Line A over a 5-minute window: use the Grafana data to estimate Availability (Productive seconds ÷ 300). Set Performance and Quality to 1.0 for now. Report the OEE percentage, then explain what fields you would add to producer.py to make Performance and Quality measurable as well.

6. Computerized Maintenance Management Systems (CMMS)

📘 Theory (Deeper Concepts)

A Computerized Maintenance Management System (CMMS) is specialized software that tracks, schedules, and manages all maintenance activities for equipment and facilities. While a MES focuses on production execution, a CMMS focuses on keeping the machines running by orchestrating preventive maintenance, reactive repairs, spare parts inventory, and maintenance history. It is the bridge between operations and reliability engineering.

Core Functions of a CMMS

  • Work order management — creates, assigns, and tracks maintenance tasks from request to completion.
  • Preventive maintenance scheduling — schedules recurring maintenance (oil changes, calibrations, inspections) based on time or machine run hours to prevent failures before they happen.
  • Predictive maintenance integration — ingests sensor data and alerts (temperature spikes, vibration anomalies) from the MES or historian to trigger work orders before equipment fails.
  • Asset management — maintains an inventory of all equipment, their specifications, maintenance history, and genealogy.
  • Spare parts tracking — manages stock levels, reorder points, and usage history to ensure critical parts are always available.
  • Labor and resource planning — assigns maintenance staff, estimates task duration, and tracks labor costs per equipment.
  • Root-cause analysis & reporting — documents failure modes, downtime duration, and lessons learned to drive continuous improvement.

Why CMMS Matters in Manufacturing

Unplanned downtime is the enemy of OEE. A single unexpected machine failure can halt an entire production line and cost thousands per hour. A CMMS minimizes this by shifting maintenance from reactive ("the machine broke, fix it now!") to preventive ("the machine is due for maintenance Tuesday night during the scheduled stop") or even predictive ("vibration sensors show early bearing wear; schedule maintenance before failure"). This shift improves equipment availability, extends asset life, and reduces emergency repair costs.

In the ball-pen example, SEMI E10 Scheduled Downtime is maintenance initiated by a CMMS work order. When the system knows in advance that maintenance is coming, it can pre-position materials, notify the scheduler, and avoid unexpected failures that would appear as Unscheduled Downtime — the enemy of OEE.

Popular CMMS Solutions

Commercial platforms include Maximo (IBM, enterprise-scale), SAP Plant Maintenance (PM), Infor EAM, Aspen Mtx, and Computerized Relative Allocation of Facilities (CRAFT, used in hospitals and facilities). Mid-market and discrete manufacturers often use UpKeep, Fiix (now part of IBM), or Dude Solutions. Open-source options like MINTO and OpenMAINT support smaller operations.

For more details, see OSApiens CMMS, which offers a comprehensive overview of CMMS platforms, best practices, and ROI calculations.

💻 Exercises (Deeper Focus)

  1. Distinguish between preventive, predictive, and reactive maintenance. Give one example of each in the ball-pen line (e.g., "replace QC camera filter every 1000 parts" is preventive).
  2. The MES and CMMS are separate systems. Explain one piece of data that must flow from the MES (e.g., parts_total run hours) to the CMMS (e.g., to trigger a PM task based on cycle count) and why they cannot live in one database.
  3. In our producer.py, add a field to record whether the line is in Scheduled Downtime (CMMS-initiated maintenance). Propose how you would trigger a scheduled downtime window and how the OEE calculation should treat that state versus Unscheduled Downtime.
  4. Look up the concept of MTBF (Mean Time Between Failures) and MTTR (Mean Time To Repair). Show how a CMMS historical report would calculate these metrics and why improving MTBF through preventive maintenance yields higher OEE than improving MTTR alone.

7. Automated Quality Control Using Vision Systems and Artificial Intelligence

📘 Theory (Deeper Concepts)

In the ball-pen simulation, quality control is rule-based: each component has a Quality flag (OK or DEFECTIVE), set randomly. In a real factory, quality is measured by inspection. Historically, this meant a human operator with a magnifying glass checking for scratches, misalignment, or dimensional errors — slow, inconsistent, and prone to fatigue. Computer vision and artificial intelligence have transformed QC into a high-speed, consistent, and measurable process that runs 24/7 without tiring.

How Computer Vision and AI Work in Manufacturing QC

Computer vision is the ability of a machine to interpret images and video. A camera on the production line captures a photo of each component or assembly. Machine learning algorithms (often deep neural networks trained on thousands of examples) instantly classify the image as pass or fail, detect specific defects (cracks, discoloration, missing features), and measure dimensions with micron precision. Unlike traditional rule-based thresholds, AI models learn subtle patterns that no engineer could codify by hand.

  • Defect detection — AI flags cracks, dents, stains, missing components, or misalignment in real time. The rejected item is removed automatically (by robot or diverter valve) before it reaches assembly.
  • Dimensional metrology — 3D cameras or structured light systems measure part dimensions in milliseconds, comparing against CAD limits. No caliper, no subjectivity.
  • Color and surface inspection — hyperspectral cameras detect coating thickness, color consistency, and surface finish anomalies that the human eye would miss.
  • Traceability & genealogy — every inspected part is photographed with a timestamp and linked to the batch, operator shift, and machine that made it. A defect pattern can be traced back instantly.
  • Continuous learning — as new failure modes appear, the model is retrained with fresh examples, making QC smarter over time.

Impact on OEE and Production

In our ball-pen line, the QualityControl station randomly rejects ~5% of inputs. With vision-based AI QC:

  • Quality improves — recall the defect rate that reaches customers drops dramatically (from 0.5% to 0.01%) because AI catches defects that humans would miss.
  • Performance stays constant — AI inspection is fast; unlike manual QC that slows the line, vision QC runs in the time it takes to snap a photo (milliseconds).
  • Availability improves — fewer escaped defects means fewer customer complaints, recalls, and warranty claims. Reputation and repeat business stay intact.

Challenges and Considerations

AI QC is not magic. Challenges include: lighting and image quality must be consistent (bad photos train bad models), the model must be retrained when product design changes, rare defects are hard to learn (you need many examples), and false positives can trigger unnecessary rejections that waste parts. The solution is human-in-the-loop: controversial images are flagged for human review, with the decision fed back into model retraining. This hybrid approach combines AI speed with human judgment.

Real-World Examples

Electronics manufacturing uses AI to inspect solder joints, component placement, and circuit board traces. Automotive plants inspect painted surfaces, weld integrity, and final assembly fit. Food & beverage uses vision to sort defective products, detect foreign objects, and verify label placement. Pharmaceuticals employ vision to count pills, check seal integrity, and detect vial cracks.

For deeper insight into computer vision in manufacturing, see Tupl: Computer Vision in Manufacturing, which covers real implementations, ROI, and case studies.

💻 Exercises (Deeper Focus)

  1. In the ball-pen QualityControl.process() method, replace the random rejection logic with a deterministic rule: "reject if the component is from a batch older than 30 days." How would you track batch age in the Component dataclass to support this? What historical data from an AI vision system would make this rule possible?
  2. Describe the training data needed to build an AI model that detects scratched barrel surfaces on ink pens. How many images are "enough"? What corner cases (lighting, orientations, magnification) would a good dataset include? Why is data labeling the bottleneck?
  3. A vision QC system rejects 3% of parts. Of those rejections, 95% are truly defective but 5% are false positives (good parts wrongly rejected). Estimate the financial cost per 1,000 pens produced if each pen costs $0.50 to make, sells for $2.00, and false rejections are scrap. Then propose one way to reduce false positives.
  4. In a real plant, the AI vision model is retrained monthly as defect patterns evolve. Propose a data pipeline: where does the training data come from (hint: human-reviewed rejects from last month), how is it labeled, and where does the new model get deployed? What test would you run before trusting the new model on the live line?
  5. Compare the OEE impact of improving Quality from 95% to 98% versus improving Availability from 80% to 83%. Which has more business impact? How does an AI vision system help you achieve each?