Skip to content
Course contents
Why live results disappoint

The ways a bot blows up

A running program can lose money faster than any human, and this chapter catalogues how: the runaway loop that fires orders in a storm, the disconnect that strands an open position, the duplicate order from a careless retry, the stale-data trade, the reaction to a bad tick. Each failure is shown with the control that prevents it.

9 min readChapter 23 of 28
What you will learn
  • Catalogue the common automated-trading failure modes
  • Map each failure to the control that prevents it
  • Internalise that a bot can lose faster than a human and needs hard limits

The great danger of automation is not that a program makes the wrong trade. It is that a program can make the wrong trade thousands of times before you look up from your coffee. A human error is one bad trade; a software error can be a bad trade repeated at machine speed until the account is gone. This chapter is a catalogue of the ways automated systems actually blow up, each paired with the control that stops it, because knowing the failure is how you build the guard.

The failure catalogue

A running program can lose money faster than any human: the runaway loop, the disconnect, the duplicate order, the stale-data trade, and the bad tick.
A running program can lose money faster than any human: the runaway loop, the disconnect, the duplicate order, the stale-data trade, and the bad tick.

The blow-ups are not exotic. They are the same few failures, again and again.

The runaway loop. A bug makes the program place orders in a tight loop, firing hundreds in seconds, each one real. This is the fastest way to empty an account, and it is guarded by hard limits: a cap on orders per minute, a maximum position, a kill switch that halts everything. Here is a runaway loop meeting the simplest such guard.

ExampleA runaway loop, stopped by a hard order-count limitch23/runaway.py
# A cautionary tale: a buggy strategy that fires an order on EVERY tick (a runaway
# loop), and the crude guard that stops it. Without the guard it would place an
# order for every one of the 50 ticks in seconds. With a simple order-count limit,
# the damage stops at the limit. This is why the next part builds real risk
# controls; here is the smallest possible version of one.
from paper_broker import PaperBroker

broker = PaperBroker(cash=1_000_000, prices={"RELIANCE": 1400})

MAX_ORDERS = 5            # the guard: never place more than this many
orders_placed = 0
halted = False

# 50 ticks arrive; a bug makes the strategy want to buy on every single one.
for tick in range(50):
    if orders_placed >= MAX_ORDERS:
        if not halted:
            print(f"KILL SWITCH: order limit of {MAX_ORDERS} reached, halting.")
            halted = True
        continue
    broker.place_order("RELIANCE", "BUY", 1, "MARKET")
    orders_placed += 1
    print(f"tick {tick}: order placed (total {orders_placed})")

print(f"\nWithout the guard, the bug would have placed 50 orders.")
print(f"The guard stopped it at {orders_placed}.")
print(f"Position: {broker.get_positions()}")
Output
tick 0: order placed (total 1)
tick 1: order placed (total 2)
tick 2: order placed (total 3)
tick 3: order placed (total 4)
tick 4: order placed (total 5)
KILL SWITCH: order limit of 5 reached, halting.

Without the guard, the bug would have placed 50 orders.
The guard stopped it at 5.
Position: [{'symbol': 'RELIANCE', 'quantity': 5, 'avg_price': 1400.0}]

The buggy strategy wanted to buy on every one of 50 ticks. Without a guard it would have placed all 50. The crude limit stopped it dead at 5 and halted, capping the damage. A real system's limits are more careful than this, but the idea is exactly this blunt: when the count crosses a line, stop, no matter what the strategy wants.

The other failures each have their own control.

  • The disconnect. Your connection drops while you hold a position, and the program is blind, unable to see the price or manage the trade. The control is to detect the disconnect and to have decided in advance what to do when blind, often to flatten and wait rather than hold something you cannot see.
  • The duplicate order. A reply is lost, the program retries, and the order goes out twice. The control is the idempotent order manager from Part 3, with its client ids.
  • The stale-data trade. The feed silently freezes, and the program keeps trading on a price that is minutes old. The control is to check data freshness and stop if it is stale, the reconciliation and monitoring habits.
  • The bad tick. A single corrupt price arrives, a zero or an absurd spike, and a naive program reacts to it violently. The control is a sanity check on incoming data: reject a price that has moved impossibly far, and do not trade on it.

The pattern behind every control

Notice what every one of these controls has in common. Each assumes the program will do something wrong and puts a hard limit in the way, outside the strategy, that cannot be reasoned around. The strategy is where the ideas live, and ideas can be buggy; the controls are where the limits live, and they exist precisely so that a buggy idea cannot cause unlimited harm. This is why the architecture keeps the risk gate separate from the strategy, and why the final part is devoted to building these controls properly. A program that can lose money faster than you can react must have limits that act faster than you can too.

What to carry forward

Automation's danger is speed: a bug can repeat a losing action thousands of times before you look up. The blow-ups are a short, familiar list, the runaway loop, the disconnect, the duplicate order, the stale-data trade, the bad tick, and each has a matching control, from hard order and position limits and a kill switch to the idempotent order manager, freshness checks, and price sanity checks. You saw the simplest of them stop a runaway loop at 5 orders instead of 50. Every control shares one shape: a hard limit outside the strategy that a bug cannot reason around. Building those limits properly, the risk gate and the kill switch, is the whole of the final part, which begins now.