20 December 2024 / Operational software

Dates, dashboards and the work that has not processed

A dashboard can look current while timezone assumptions or hidden processing limits tell a different story. I am putting unprocessed records, date rules and result limits into the interface because operational users need to see incomplete work rather than discover it through a customer complaint.

Intersecting steel beams and bracing viewed from beneath a bridge.
Photo: Sebastian Schuster (opens in a new tab)

"Today", "complete" and "all records" each depend on a rule. If the dashboard hides those rules, its query may be internally consistent while the result is wrong for the person using it. Operational screens should expose the boundaries behind the number. That includes timezone, data freshness, processing state, filters and any limit on the rows considered. A timestamp represents an instant. A calendar day belongs to a timezone. Grouping records by date without choosing that timezone moves work around the boundary, which is especially visible around midnight. Store event timestamps in a consistent instant format such as UTC, while retaining the source timezone when it carries business meaning. Convert to the reporting timezone before deriving a local date. Do not derive a UTC date first and then relabel it as local. Choose whether the dashboard follows the organisation, site, user or underlying event. A national operation may use one reporting timezone for financial totals and local site time for appointments. Put that choice in the metric definition rather than letting the browser decide silently. Display the active timezone near date filters and summaries. Use an unambiguous label such as 20 Dec, 9:00 am AEDT when the exact instant matters. Avoid abbreviations alone when users work across regions, because some abbreviations are reused.

Test records immediately before and after midnight in every supported reporting zone. Include daylight-saving transitions where they apply. Many records have more than one relevant time. An order may be created, received by an integration, processed by a worker and displayed later. Choosing the wrong one can make a backlog disappear from one day and appear in another. Name timestamps for the events they represent: created_at, received_at, processing_started_at, processed_at and failed_at. A generic date column invites several parts of the application to assign different meanings. Use event time when answering when the business event happened. Use processing time when measuring the pipeline. If a dashboard says "received today", it should not filter on the day a delayed job finally completed. Late-arriving records need a policy. Historical business totals may change when delayed events are accepted, while a daily processing report should show that they arrived late. Keep both facts rather than overwriting the event time with the import time. The screen should state which timestamp drives the filter, particularly when users can switch between created, due, processed and settled dates. A completed-work chart cannot reveal items that never reached the completed table or status. Put backlog and failure counts near the result they qualify. Distinguish records that are waiting normally, currently processing, delayed beyond an expected threshold, rejected for correction and stuck after a worker interruption. A single "not processed" count hides whether staff should wait, fix data or restart a job.

Age matters. Ten new items waiting for the next scheduled run differ from ten items that have waited for two days. Show the oldest waiting time and useful age bands. Link counts to a filtered list so staff can inspect the actual records. Also count work that failed before creating the usual domain record. Integration inboxes, upload staging tables and dead-letter queues may contain items absent from the main application. The dashboard needs a source for those pre-processing failures or an explicit statement that it does not include them. Do not subtract failed work from the denominator and then call the remainder a success rate. Dashboards often protect performance by limiting rows, date range, API pages or chart points. If the interface does not reveal that limit, users naturally read the displayed result as complete. Apply filters and aggregation in the data store before limiting output wherever possible. Fetching the first thousand rows and then counting them is not a count of all matching records. If an upstream API provides pages, follow pagination until the defined scope is complete or label the result partial. When a limit is intentional, show it. "Displaying 500 of 2,346 matching records" gives the user a decision. If the total is unknown, say that the result was truncated and offer a narrower filter or background export.

Charts can aggregate large result sets without rendering every row, but the aggregation query still needs a bounded and documented scope. Sampling may be valid for exploratory analysis, not for a queue that staff expect to clear. Test just below, at and above every configured limit. Silent truncation often survives ordinary test data because the fixture never becomes large enough. A dashboard may render now using a cache, replica, warehouse or scheduled extract that last changed earlier. Display the source update time separately from the page load time. Show the last successful data update and, where useful, the oldest source event not yet processed. Track the health of the update itself. If a scheduled refresh fails, retain the last good data but mark it stale rather than displaying it as current. Freshness needs a threshold tied to the workflow. A monthly report and a live support queue have different tolerances. Once the threshold is exceeded, change the visual state and notify the owner responsible for restoring the pipeline. Avoid using the newest record timestamp as the only freshness signal. A quiet period can look stale, while one newly processed record can hide a large backlog. Combine update health with queue state.

Include source and refresh metadata in exports so a downloaded file can still be interpreted later. A useful acceptance set starts at the edges: midnight, daylight-saving changes, delayed events, failed work, empty periods and record counts beyond the display limit. For each metric, write down the source records, included statuses, driving timestamp, reporting timezone, refresh path and row-limit behaviour. Then create a small fixture where the expected count can be calculated by hand. Compare the summary, chart, filtered detail and export against the same definition. Add reconciliation checks between stages of the pipeline. Received should equal processed plus waiting plus failed, within any explicitly excluded states. Alert on a broken equality rather than waiting for somebody to notice that a chart looks light. For the release check, ask an operational user to explain one number from the screen: which records it includes, as at what time, in which timezone and what remains outside it. Note anything they had to ask a developer or inspect source code to discover, then add that information to the metric definition or interface.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs