The node ships to a rooftop cabinet you cannot easily reach. Before it goes out: right-size it, lock it down, and make it tell you when it is sick. You will shrink an image, get a leaked credential out of it, harden two containers without breaking the GPU, add health checks and auto-restart, and write an edge-local alerter that reports incidents instead of noise.

HW05 gave Ops a benchmark and a recommended config. Ops signed off, and now the node has to survive unattended on a roof. Nobody is going to SSH in when it wedges at 3am, the cabinet is physically reachable by strangers, and the device is shared. A service that works on your desk and a service you can deploy are different things; this homework is the difference.

Full step-by-step instructions, including every command you need to run, are in the README.md of the repository Lumen creates for you. Work from that README. This page is the summary and the requirements.


Objective and Expected Learning Outcomes

By completing this assignment, you will be able to:

  1. Choose a runtime configuration from measured trade-offs, re-checking accuracy and not just speed.
  2. Write a multi-stage Dockerfile and explain where the savings came from.
  3. Keep secrets out of image layers and out of Git, and prove it.
  4. Harden a container (non-root, read-only, dropped capabilities, no new privileges, cgroup limits) without breaking the workload, including a GPU workload.
  5. Add health checks and restart policies that recover a service automatically.
  6. Build edge-local alerting that debounces, so it survives contact with a real operator.

The Device

All work runs on the shared course DGX Spark at cs494.evl.uic.edu, over SSH. Because it is shared: prefix everything with ${USER}, keep files under your home directory, and remember that resource limits are not bureaucracy. An unbounded container starves your classmates.

You can harden only your own account and containers. Device-wide settings such as the SSH daemon, firewall, and OS updates belong to your instructor, and you only inspect those.

“Root” means two different things here. The node runs rootless Podman, so a container process reporting uid=0 is already mapped to your account on the host, not to system root. Hardening is still worth doing, but be precise about what it buys: dropping privileges limits what a compromised process can do inside its own namespace and to the files you mounted, rather than standing between an attacker and the machine. That second wall is already there. Say which of the two you mean when you defend a control in your report.

Take the stack down when you are done for the session. On a shared device, leaving containers running is rude.


Instructions

Step 1: Bring the Stack Up As-Shipped

SSH to the device, export the instructor’s base-image tag and your host user and group ids, and bring the stack up once to see it work before you change anything. It runs as root inside the container, unrestricted, with the alerter unimplemented. That is your starting point, not your deliverable.

Step 2: Optimize by Picking the Runtime Config

The cheapest optimization is the one you already proved in HW05: the smallest model and input size that is still accurate enough. Two knobs go further: input size and half precision on the GPU. Run the provided comparison script, which reports latency and detections per frame. Read both columns: a config is only better if it is faster and its detections have not collapsed. A smaller input size misses small and distant vehicles; at an intersection, that is a missed car, not a rounding error. Optimization is only valid if you re-measure accuracy, not just speed.

Step 3: Slim the Image

The provided fat Dockerfile works and is awful. Read it and count what is wrong with it: size problems, security problems, and one line that is both. Then write a multi-stage replacement and complete the two missing ignore-file entries. Re-measure after each change. The savings are not evenly split across the steps; your report has to say which single change bought the most, which means measuring rather than guessing. Then confirm the image still works.

Step 4: Get the Secret Out

The fat image bakes in an ops-center token. Prove the leak first; you cannot argue you fixed something you never demonstrated. Then remediate: create your run-time secret file from the committed template, lock its permissions down, make sure Git can never see it, and wire it into the alerter service at run time. Verify all three: absent from image history, ignored by Git, and absent from status. The repository’s ignore file does not already cover that filename. Adding the rule is a real step; check, do not assume.

Step 5: Harden Both Services

Complete the hardening settings in the Compose file: non-root user, read-only root filesystem plus a writable tmpfs, all capabilities dropped, no new privileges, and memory and CPU limits on both services. Bring the stack back up and collect the evidence with the provided script, which reads the running containers: the user they run as, the cgroup limits the kernel is enforcing, whether the root filesystem really is read-only, and whether CUDA is still available. Claiming a setting in YAML is not the same as it taking effect.

This is the hard part: hardening that breaks the service is not hardening. If one specific setting cuts the vision container off from the GPU, that is a finding. Relax that one, keep the rest, and say exactly which and why in your report. “I dropped everything and it worked” and “I had to keep exactly one” are both correct answers. “I gave up on hardening” is not.

Step 6: Health Checks and Auto-Restart

Add a health check and a restart policy to each service. The probe is provided and exits non-zero when a service’s heartbeat file stops being refreshed. It is written in Python because a minimal image often ships neither curl nor wget, and it watches a heartbeat rather than a port because a wedged loop keeps its port open. Verify health, then verify recovery by killing a service and confirming it comes back on its own. Right after startup a service may briefly report as starting; that is what the start period is for. Too short a start period reports a healthy service as failed.

Step 7: Wire the Threshold Alert

Complete the alerter. It tails the vision service’s telemetry and decides on the device that something is wrong, with no cloud round trip. Implement which rules a record breaches, remembering that records can be incomplete where no thermal zone is readable; the debounced raised-and-cleared state machine, which is the graded piece; and the heartbeat write for the health check.

Tune the alert limits against your own measured baseline from the device. The shipped thermal value is a generic starting point, not a measurement. Measure your own idle and loaded baselines before choosing a number, and say in the report which one you measured. A threshold nobody can defend is the same as no threshold: one that never fires and one that fires constantly are equally useless, and the second one gets your alerter muted by the operator it was supposed to protect.

Step 8: Write the Hardening Report and Reflection

Cite your numbers and your evidence files in the report template, then fill in the reflection with real measurements.


Submission Requirements

To receive credit, you must:

  1. Work on the development branch in your Lumen-provisioned repository.
  2. Commit your slimmed Dockerfile, completed ignore file, and Compose file.
  3. Commit your completed alerter script and tuned alert limits.
  4. Commit the four evidence files: config comparison, image sizes, hardening evidence, and real alerts.
  5. Commit your filled-in hardening report and reflection.
  6. Ensure evidence comes from your own run on the device. Simulate mode exists so you can develop the alerting logic off-device; simulated records are stamped as simulated and are not submittable evidence.
  7. Push to development and open a Student PR from development into main.

Do not merge the Student PR. Never commit a real secret file: the real token file must be git-ignored, while the template is committed. Do not commit videos, model weights, or .venv.


Evaluation Criteria

Criterion Weight
TODO 1: multi-stage image that is meaningfully smaller and still runs; .dockerignore completed 20%
TODO 2: secret out of the image layers and out of Git, with proof of both 15%
TODO 3: both services hardened, verified from hardeningEvidence.txt: the telemetry file on the host is owned by you, the container cannot escalate, and the GPU is still reachable (or the exception named and justified) 25%
TODO 4: health checks + restart policy, with recovery demonstrated 10%
TODO 5: debounced alerter: incidents, not per-sample noise; tolerates missing fields 20%
hardeningReport.md + reflection: measured numbers, defended thresholds, accuracy re-checked 10%

Hardening is graded on the evidence from your running containers, not on what your compose file claims.


Notes on Course-Wide Requirements

  • A reflection file is required for this assignment, including the certification statement in its header.
  • The workflow is unchanged: development branch, Student PR into main, and no merging of your own PR.
  • This assignment pairs with Lab 07 (Optimization, Monitoring and Security). The lab ran each technique once on a throwaway container; here you apply all of them to a real two-service stack that has to keep running afterwards, which is where they start to conflict with each other. Container and GPU mechanics come from Lab 02 Parts 14 through 18.
  • The vision service is the HW04 and HW05 counting pipeline turned into a long-running service, shipped complete so this homework never depends on your own earlier work. You do not edit it; you deploy it properly.
  • Your instructor verifies the shared device each semester. If a step fails in a way that looks like the device rather than your code, ask them rather than trying to debug a machine the whole class shares.

Additional Resources