Course contents
Watching a system that never sleeps
You cannot watch a live system every second, so it must tell you what it is doing and shout when something is wrong. This chapter covers logging every decision and order, monitoring live positions and profit and loss, and alerting on the anomalies that need a human now: a breached limit, a disconnect, an unexpected position.
- Log every decision, order, and fill in a form you can audit later
- Monitor live positions and profit and loss against expectations
- Define alert conditions that demand immediate human attention
A live trading program works while you sleep, while you are at your day job, while you are not looking. That is the point of it, and it is also the danger. A system you are not watching can drift into trouble unseen. The answer is not to watch it every second, which is impossible, but to make it tell you what it is doing and, above all, to make it shout when something needs you now. That is logging, monitoring, and alerting.
Log everything
A log is a written record of what the program did, line by line, as it did it. Every decision, every order placed, every fill received, every error, with the time it happened. Logging feels tedious to write and is priceless the first time something goes wrong, because the log is the only account of what actually occurred. When a position appears that you did not expect, or the day ends worse than it should have, the log is where you find out why. Log enough that you could reconstruct the whole day from the log alone. Here is the shape of it, with a small monitor attached.
# Logging and monitoring. A live system logs every action so you can see later
# what it did, and it raises an ALERT (a warning) the moment something needs a
# human now. The log lines here omit the timestamp so the output is stable; a
# real system stamps every line with the time.
import logging
logging.basicConfig(level=logging.INFO, format="%(levelname)-7s %(message)s")
log = logging.getLogger("bot")
MAX_POSITION = 200
PNL_ALERT = -15_000 # alert if the running loss reaches this
def monitor(position, pnl):
if abs(position) > MAX_POSITION:
log.warning(f"ALERT: position {position} over limit {MAX_POSITION}, look now")
if pnl <= PNL_ALERT:
log.warning(f"ALERT: pnl {pnl} at or past limit {PNL_ALERT}, look now")
# A short run. Each step is logged, then checked by the monitor.
steps = [
("bought 100 RELIANCE at 1400", 100, 0),
("price 1380, marked to market", 100, -2000),
("bought 150 more (position now 250)", 250, -2000), # over the position limit
("price 1250, marked to market", 250, -34500), # a large loss
]
for description, position, pnl in steps:
log.info(f"{description} | position {position}, pnl {pnl}")
monitor(position, pnl)
log.info("run complete; review the ALERT lines above")The run logs each step as it happens, an ordinary INFO line, the routine record. But watch the two louder lines. When the position crosses the limit of 200, and when the loss passes 15,000, the monitor raises an ALERT, a warning that stands out from the routine log precisely because it needs a human to look now. That is the distinction that matters: most log lines are for reading later, but an alert is for acting on immediately.
Monitor against expectations
Monitoring is checking, continuously, that the system's real state matches what it should be. Your live position, your profit and loss for the day, the number of open orders, the time since the last tick arrived: each has an expected range, and monitoring watches for a step outside it. The power is in comparing to expectations. A position of 250 is not wrong in the abstract; it is wrong because your limit was 200. A quiet feed is not wrong until it has been quiet longer than a live market ever is. Good monitoring encodes what normal looks like and flags the moment reality leaves it.
Alert on what needs you now
The final piece is deciding what deserves to interrupt you. An alert should fire for the things a human must handle at once and nothing else, or you will learn to ignore it. A breached risk limit. A disconnection from the broker. A position that does not match what the strategy believes it holds. The day's loss approaching its cap. Each of these is a case where the program has reached the edge of what it can safely handle alone and needs a person. Send those alerts somewhere you will actually see them, and keep the routine noise in the log where it belongs. A system that cries wolf is one whose real warnings get missed.
What to carry forward
Because a live system runs while you are not watching, it has to tell you what it is doing and shout when something is wrong. Log every decision, order, and fill with the time, so any day can be reconstructed from the log alone. Monitor the live position and profit and loss against the ranges you expect, and flag anything outside them, as the monitor flagged the position breaching its limit and the loss passing its cap. Reserve alerts for what needs a human at once, and keep everything else in the log, so a real warning is never buried. Next, the discipline that lets you trust the code doing all of this in the first place: testing.