Skip to content

Monitoring, history and alerts

One grid covers every node, updating every second. Each row is one study moving through one node in one direction: an association delivering it, or the node sending it to one destination. Columns: node, AE titles, source and route (or destination), accession, study, modality, instance counts (sent, queued, failed), size and the last error. Senders turned away in the last hour are listed below it.

The Analytics tab charts how things have been going, where Monitoring shows what is happening now.

Nodes and queues, over the last hour, 6 hours, 24 hours or 7 days:

  • Instances per minute: received and sent by all nodes, and dead-lettered.
  • Waiting to be sent, per destination, all nodes together. This shows a destination falling behind long before its queue raises an alert.
  • Longest wait: the oldest instance waiting for each destination, and which destinations were unreachable at times.
  • Disk space free on each node, and, when it is going down, when it would run out at that rate.

The central service samples every node once a minute and keeps the samples for 7 days.

Usage, counted from the history:

  • Per hour today, against the same day last week. You can split it by modality, node or destination.
  • When instances arrive: the average per weekday and hour over four weeks. Use it to choose when destination schedules pause sending, or when to limit bandwidth.
  • Studies per day and problems per day (instances refused, deliveries failed) over 30 days.
  • Destinations: delivered, failed, pending, success rate, average time to deliver and the time 95% were delivered within, over 30 days.
  • Senders refused over 30 days: usually a modality whose AE title no route accepts.

Hourly counts are kept 400 days, daily counts for good, both after the history itself is deleted. The charts contain no patient data, so everyone who can use the console sees them.

The Nodes tab shows each node’s state, listening ports, associations, queues and rates, received and forwarded counts, and disk space. Expand a node for each destination’s queue (oldest instance, unreachable, paused, rate-limited), last error and dead letters, with buttons to re-queue or purge them. Drain and Resume are here too.

History searches every instance the nodes received, held, filtered or refused: by accession, patient ID, patient name (from its start: SMITH^J), study or instance UID, sender, time, or problems only. Each study shows its deliveries per destination, and each instance its state at every destination, with errors. It stays fast on large histories.

  • Resend a study to any destination or group: each node sends again the copies it holds, from its resend cache (24 hours by default), queues or dead letters.
  • Open in viewer: with Settings › Viewer address set (for example https://viewer.example.org/viewer?StudyInstanceUIDs={StudyInstanceUID}, or {AccessionNumber} for viewers that look studies up by accession), each study has a link that opens it in that web viewer. An OHIF viewer can read the study through the nodes’ DICOMweb: see Connect an OHIF viewer.
  • History is kept 30 days by default (Settings), and nodes keep it on disk while the central service is unreachable.

Other views on the History tab: Dead letters, Quarantine, Prior studies, AI jobs, Reconciliation, HL7 messages and Worklist.

  • Dead letters: instances a destination refused for good, or that used every attempt. Re-queue them after fixing the cause, send them elsewhere, purge them, or download one.
  • Quarantine (off by default; Settings › Quarantine (days)): refused instances kept on the node, with the reason, for diagnosis. After fixing the source or route, Route again sends them through routing as if just received.

The history contains patient names and IDs: only administrators and history viewers see it, and every search is audited.

Every 15 seconds the central service raises alerts, and resolves them when the condition clears:

  • a node offline (except nodes drained on purpose), or a node problem lasting over a minute;
  • a destination unreachable; a queue too deep or its oldest instance too old; dead letters present;
  • low disk space;
  • a sender turned away;
  • a central server stopped;
  • an HL7 queue stopped or not getting through, MPPS or status updates waiting, AI results overdue, studies waiting for reconciliation.

Thresholds are on the Alerts tab. Notifications go by email (SMTP) and/or a webhook (Teams and Slack incoming webhooks work as they are), when an alert is raised or resolved, with optional reminders.

The Reports tab counts usage over any range of dates, per day, week or month: per sender, source, modality, route or node (instances, studies, volume, refused), and per destination (delivered, failed, pending, average time to deliver, and the time 95% of deliveries took at most: a slow tail the average hides), with a chart, a table and CSV. Scheduled emails send a daily, weekly or monthly report. Reports contain no patient data.

  • Prometheus: /metrics on any central server exposes every node and destination: state, counters, queues, disk, open alerts. Turn it on under Configuration › Node enrollment › Metrics, which makes a token.
  • Logs: each service writes daily log files, kept 30 days. Settings › Logging also sends them to the Windows event log and to a syslog server or SIEM (UDP, TCP or TLS). Logs can name patients: send them only where patient data may go.
  • OpenTelemetry: metrics, traces and logs over OTLP (below).

With Settings › OpenTelemetry › OTLP endpoint set, the central servers and nodes send what they know over OTLP, the OpenTelemetry protocol, to a collector or straight to a monitoring service that accepts it (Grafana, Datadog, Splunk, Dynatrace, New Relic, Elastic, Azure Monitor and others). Headers, often an API key, are kept like a password.

  • Metrics: the same as the Prometheus endpoint (routes_node_queued_instances, routes_destination_unreachable and the rest, counters without _total), sent every 30 seconds by the leading central server only, so that they are not counted twice; and each service’s runtime metrics (memory, garbage collection, threads), and request durations by route on the central servers.

  • Traces: the path of each instance through a node, and why it took as long as it did:

    Span From, to Attributes
    dicom.receive the header arriving, to the routing decision calling and called AE, sender’s address, source, route, SOP class, modality, transfer syntax, size, outcome (queued, held, held for reconciliation…) and destinations
    dicom.deliver queued, to delivered, dead-lettered, discarded or moved to another member of a group destination, its AE, host and port, attempts, size, transfer syntax sent, outcome; each retry is an event with its reason
    routes.study.release the first instance of a held study, to its release route, source, sender, instances, outcome

    Deliveries are children of their receipt, so a trace shows the whole wait: a destination down for 40 minutes is a dicom.deliver span 40 minutes long with its retries in it. Console requests on the central servers are traced too, named by their route (GET /api/history); nodes’ calls, the live feed and health checks are not.

  • Logs: off by default. When chosen, each line goes with the trace it was written in.

Nothing that names a patient goes in a span: no patient IDs, names, accession numbers or study and instance UIDs; requests are named by their route, not their path or query, and the addresses the central servers call (a webhook’s can hold its token) are cut to their host. Log lines can still name senders and UIDs, which is why logs are off unless chosen. Share of traces sends fewer for busy sites: each trace is kept or dropped whole.

Metrics need the central servers to hear from the nodes (as the Prometheus endpoint does); traces and logs go from each node straight to the collector, so nodes must be able to reach it.