The pump station is running, benchmarked and hardened. Now it has to talk to the operations center. You will carry three different kinds of message over three different protocols, measure what each one costs, prove with an experiment what happens to telemetry when the ops center goes offline, and recommend which transport belongs on which link.

The city’s operations center wants three things from your node, and they are not the same kind of traffic:

They want Rate Who starts it If it is late
Telemetry: every sensor reading high, continuous the node a lost sample is survivable
Summaries: “what did the pump do this hour?” rare, on request the ops center nothing breaks
Alerts: “the pump is failing NOW” rare, urgent the node somebody replaces a bearing instead of a pump

One protocol for all three is a design mistake in one direction or another. Your job is to build all three links, measure them, and defend the split.

Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.


Objective and Expected Learning Outcomes

By completing this assignment, you will be able to:

  1. Publish and subscribe telemetry through an MQTT broker, and measure what was delivered, not just assume it arrived.
  2. Serve a request/response API for on-demand queries.
  3. Push live events to connected clients over a WebSocket.
  4. Measure latency, connection cost and throughput per protocol, and read the results honestly.
  5. Explain MQTT delivery guarantees, QoS and session persistence, and demonstrate the difference experimentally.
  6. Match a protocol to a message class and defend the choice with data.

The Device

All work runs on the shared course DGX Spark at cs494.evl.uic.edu, over SSH. This assignment binds three network ports, so it matters more than usual that you use only your own three ports, which you derive from your UID, and that every container resource and MQTT topic you create is prefixed with ${USER}. Two students on one port is an afternoon lost to debugging each other’s telemetry.

Clean up after yourself on the shared device: bring the stack down when you are done.


Instructions

Publish each telemetry record to the node’s topic, noting when you re-stamp the sent timestamp: the receiver is timing the transport, not how long the record sat in the generator. Then complete the subscriber: the client identity and session settings the delivery experiment turns on, the measurement itself (latency from the sent timestamp, loss and duplicates from the sequence number, and appending each record to the shared bus), and the subscribe call placed on connect rather than once at startup. Think about what happens on a reconnect if you do it the other way.

The subscriber writes its statistics when it stops, not while it runs, and it rewrites the file from scratch on every shutdown. Once you have a run you are happy with, copy it, and submit the copy if the live file gets clobbered.

The publisher replays a window that covers the cavitation failure event, so there is real failure data on the wire, not just noise.

Complete the three endpoints: health, latest reading, and a windowed summary. Getting the happy path working is the easy half. The graded half is what your API does when there is no data yet and when a caller sends nonsense. An HTTP client can only tell those apart if your status codes say so.

Hold each operator connection open, push to all of them concurrently and survive dead ones, and watch the telemetry bus for threshold crossings. Then copy the dashboard to your own machine, fill in the two values in its config block, and open it in a browser. Re-run the publisher and watch alerts appear with no page reload. The recording’s three failure events give a correct implementation windows to alert inside; each alert carries its timestamp and the manifest’s ground-truth flag so you can check.

The default limits are starting points measured from this recording, whose channels are normalized proxies rather than real units. Justify whatever limits you finally use.

Step 4: The Delivery-Guarantee Experiment

This is the hard part, and it is an argument you have to earn with data. The provided script publishes three batches while killing the ops-center subscriber in the middle, twice: once at QoS 0 with no session, once at QoS 1 with a persistent session. It then prints exactly which sequence numbers never arrived.

The second run changes two things at once, the QoS and the session, so the result does not tell you which one did the work. Run at least one more scenario yourself, changing only one of them, and report what you found. That one-variable-at-a-time discipline is the same rule as the HW05 sweep. Then answer, in the report, what the ops center should actually use for telemetry, knowing what each guarantee costs.

Step 5: Measure and Compare

Measure request/response cost both with a reused connection and with a new one per request, and one-way delivery latency through the real broker. The WebSocket measurement is provided complete.

Then fill in the protocol report, addressing two warnings head-on. You are comparing a one-way number against round trips; say so, and do not quietly rank them as if they were the same thing. And everything here runs on one machine over loopback, so connection setup looks nearly free. Over a real link to a cloud API, with DNS, a WAN round trip, and a TLS handshake, it is not. Reason about what changes, and be explicit that this part is reasoning, not measurement.

Cite protocol framing costs as documented costs; the harness does not measure bytes, and a number you cannot verify is worse than no number.

Step 6: Reflection

Fill in the reflection with your real measurements.


Submission Requirements

To receive credit, you must:

  1. Work on the development branch in your Lumen-provisioned repository.
  2. Commit your completed publisher, subscriber, HTTP server, WebSocket alert server, and comparison harness.
  3. Commit the MQTT statistics from a real run.
  4. Commit the protocol comparison outputs in JSON, CSV, and text form.
  5. Commit the WebSocket alerts log and the drop-test output.
  6. Commit a screenshot of your dashboard with at least one alert on it.
  7. Commit your filled-in protocol report and reflection.
  8. Push to development and open a Student PR from development into main.

Do not merge the Student PR. Use only your own derived ports, and prefix every container resource and MQTT topic with ${USER}.


Evaluation Criteria

Criterion Weight
TODO 1 + 2: MQTT publish/subscribe working, with correct latency, loss and duplicate measurement 25%
TODO 3: HTTP endpoints, including sensible behavior for no-data and bad input 15%
TODO 4: WebSocket push: alerts reach a live browser client; dead clients do not break the server 20%
TODO 5 + comparison outputs from a real run 15%
Part F: delivery experiment run, plus your own one-variable scenario, correctly explained 15%
protocolReport.md + reflection: protocol-per-message-class recommendation defended with your numbers 10%

The recommendation is graded on whether it is argued from your measurements and the delivery experiment, not from the lab’s summary table.


Notes on Course-Wide Requirements

  • A reflection file is required for this assignment, including the certification statement in its header.
  • The workflow is unchanged: development branch, Student PR into main, and no merging of your own PR.
  • This assignment pairs with Lab 08 (MQTT / HTTP / WebSocket). The lab moved one random number over each protocol; here each protocol carries the message class it is actually right for, over the real PumpWatch stream from HW02 and HW03, including its three genuine failure events. If you are stuck on a step, read its lab part first.
  • The lab’s own summary table is where most reports stop. Yours has to go further: the delivery experiment is about guarantees the lab only mentions in passing.
  • Your instructor verifies the shared device each semester. If a step fails in a way that looks like the device rather than your code, ask them rather than trying to debug a machine the whole class shares.

Additional Resources