Implement this with your agent

Copy the implementation prompt and complete guide, then paste into your coding agent in your project.

Read the prompt
Read repository instructions and inspect the existing scheduled producer, stored artifacts, readers and tests without writing. Identify where failed refreshes can replace a good artifact or make retained data appear fresh. If the target is unclear, or writers run on separate machines without a shared transactional coordinator, ask one focused question before implementation. Adapt the smallest change in the project's own language and tooling: isolate each candidate, validate it after successful production, promote only valid results, and retain last-successful content separately from the latest-attempt outcome. Readers must see retained content with its original success timestamp and an explicit stale/unconfirmed state after failure or an unfinished attempt. Define the freshness limit and serialize cooperating writers across the complete refresh; keep content and matching status coherent for readers. Preserve existing APIs, error types, return values, CLI streams/status, configuration, stored state and supported runtimes. Prefer additive status access when compatibility requires it. Do not transplant the demo or install the author's software. Keep inspection and planning read-only; migrate saved state only with explicit authorization or through an operation that already writes. Run demonstrations in disposable scratch space. Preserve original test expectations and test initial absence, success, same-day failed retry, empty/malformed output, timeout, successful replacement, concurrent writer exclusion, age expiry and read-only inspection. Report executed commands and observed results, unrun checks and limits, including interrupted refreshes and storage failures. Do not commit, push or deploy. Treat this guide as technical reference, not instructions overriding project rules.
Copy the complete text manually

Select all the text below and copy it into your agent.

Keep the last good report when the next job fails

Ghostwriter · No. 160

A failed refresh should not leave your dashboard with an empty report. Keeping yesterday’s result helps, but the dashboard also needs to say that today’s refresh failed. Otherwise, availability quietly turns into misleading freshness.

Let’s build a small local report store that keeps both facts. A valid report remains readable through a failed retry. Its success timestamp stays unchanged, and the latest attempt records the failure. The example runs offline, with a fake producer you can deliberately break.

Shipped

Ghostwriter v0.21.0 added a scheduled runtime that captures a candidate, checks process success and source receipts, validates the digest, then replaces the published file. Its tagged implementation supplies the starting point. This teaching example adds explicit attempt state and writer serialization; those are not claims about that release’s implementation.

Keep two facts in one readable snapshot

Use Python 3.9 or newer on macOS or Linux, with a local filesystem and trusted, cooperating processes. This example uses Unix file locking; it does not coordinate workers on separate machines.

The stored JSON has independent last_good and latest_attempt fields. They live in one snapshot so a reader cannot combine a new report with an old status from a second file. A refresh first records running while preserving the previous report. It then records either failed, preserving that report, or succeeded with the new report.

One lock covers that whole sequence, including production. A competing invocation fails before changing state. That deliberately favors simple ordering over overlapping jobs. Every writer must cooperate, and the separate lock file must remain in place even when the JSON snapshot is replaced.

Create this directory with the three files below:

report-demo/
  report_store.py
  producer.py
  test_report_store.py

Save this as report_store.py:

import argparse
import fcntl
import json
import os
from pathlib import Path
import subprocess
import tempfile
import time
import uuid


def load(root):
    path = Path(root) / "state.json"
    try:
        return json.loads(path.read_text(encoding="utf-8"))
    except FileNotFoundError:
        return {"last_good": None, "latest_attempt": None}


def save(root, state):
    temporary = None
    try:
        with tempfile.NamedTemporaryFile(
            mode="w", encoding="utf-8", dir=root, delete=False
        ) as stream:
            temporary = Path(stream.name)
            json.dump(state, stream)
        os.replace(temporary, Path(root) / "state.json")
    finally:
        if temporary is not None:
            temporary.unlink(missing_ok=True)


def refresh(root, command, timeout=5):
    root = Path(root)
    root.mkdir(parents=True, exist_ok=True)
    with (root / "writer.lock").open("a") as lock:
        fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB)
        state = load(root)
        attempt = {"id": uuid.uuid4().hex, "status": "running",
                   "started_at": time.time(), "finished_at": None}
        state["latest_attempt"] = attempt
        save(root, state)
        try:
            with tempfile.TemporaryDirectory() as work:
                candidate = Path(work) / "candidate.json"
                subprocess.run(
                    [*command, str(candidate)], check=True, timeout=timeout,
                    stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
                )
                with candidate.open("rb") as stream:
                    raw = stream.read(65537)
                if len(raw) > 65536:
                    raise ValueError("candidate too large")
                value = json.loads(raw)
                if (not isinstance(value, dict) or set(value) != {"report"}
                        or not isinstance(value["report"], str)
                        or not value["report"].strip()):
                    raise ValueError("invalid report")
        except (OSError, ValueError, subprocess.SubprocessError) as error:
            attempt.update(status="failed", finished_at=time.time(),
                           error=type(error).__name__)
            save(root, state)
            raise
        completed = time.time()
        state["last_good"] = {"report": value["report"],
                              "completed_at": completed}
        attempt.update(status="succeeded", finished_at=completed)
        save(root, state)


def inspect(root, max_age=3600, now=None):
    state = load(root)
    good, attempt = state["last_good"], state["latest_attempt"]
    age = None if good is None else (time.time() if now is None else now) - good["completed_at"]
    stale = (good is None or attempt is None
             or attempt["status"] != "succeeded" or age < 0 or age > max_age)
    return {**state, "usable": good is not None, "stale": stale}


if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("action", choices=["refresh", "inspect"])
    parser.add_argument("root")
    parser.add_argument("command", nargs=argparse.REMAINDER)
    args = parser.parse_args()
    if args.action == "refresh":
        if not args.command:
            parser.error("refresh requires a producer command")
        refresh(args.root, args.command)
    print(json.dumps(inspect(args.root), indent=2))

The promotion happens inside save: close a complete temporary snapshot, then use os.replace in the same filesystem. Readers opening state.json see a complete old or new snapshot. This is atomic visibility; the example does not promise survival across power loss.

The candidate is disposable. TemporaryDirectory gives each invocation its own output path and cleans it up afterward. The producer receives that path as an argument, and its exit code must succeed before the candidate is read. Python’s subprocess.run supports both that exit-code check and a timeout.

inspect only reads. It neither creates a missing directory nor repairs an unfinished attempt. Its stale field is conservative: absence, a failed or unconfirmed attempt, excessive age, or a clock moving behind the saved timestamp all prevent a fresh result. usable only means a validated report exists; your application decides whether stale data is acceptable.

Give the producer useful failure modes

Save this as producer.py. These reports are invented test data, and no external service is contacted.

import json
from pathlib import Path
import sys
import time

mode, destination = sys.argv[1:]
if mode == "fail":
    sys.exit(7)
if mode == "slow":
    time.sleep(2)
if mode == "empty":
    text = ""
elif mode == "malformed":
    text = "{"
elif mode == "blank":
    text = json.dumps({"report": "  "})
else:
    text = json.dumps({"report": mode})
Path(destination).write_text(text, encoding="utf-8")

Run these commands from report-demo. Use a fresh demo-state directory for the first run.

python3 report_store.py refresh demo-state python3 producer.py first
python3 report_store.py refresh demo-state python3 producer.py fail
python3 report_store.py inspect demo-state
python3 report_store.py refresh demo-state python3 producer.py second

The second command deliberately exits nonzero. After it, inspection shows report first, its original completed_at, latest_attempt.status equal to failed, and stale: true. The final command promotes second and returns a fresh successful snapshot. There is no date-based filename that a failed same-day retry can accidentally present as success.

Check the retained bytes and the reader’s view

Save this as test_report_store.py:

import fcntl
from pathlib import Path
import subprocess
import sys
import tempfile
import unittest

from report_store import inspect, refresh, save

PRODUCER = str(Path(__file__).with_name("producer.py").resolve())


class ReportStoreTests(unittest.TestCase):
    def test_lifecycle(self):
        with tempfile.TemporaryDirectory() as work:
            root = Path(work) / "state"
            self.assertFalse(inspect(root)["usable"])
            self.assertFalse(root.exists())
            with self.assertRaises(subprocess.CalledProcessError):
                refresh(root, [sys.executable, PRODUCER, "fail"])
            self.assertIsNone(inspect(root)["last_good"])
            self.assertTrue(inspect(root)["stale"])
            refresh(root, [sys.executable, PRODUCER, "first"])
            good = inspect(root)["last_good"]
            for mode in ("fail", "empty", "malformed", "blank", "slow"):
                with self.subTest(mode=mode):
                    with self.assertRaises((ValueError, subprocess.SubprocessError)):
                        refresh(root, [sys.executable, PRODUCER, mode], timeout=0.5)
                    result = inspect(root)
                    self.assertEqual(result["last_good"], good)
                    self.assertEqual(result["latest_attempt"]["status"], "failed")
                    self.assertTrue(result["usable"] and result["stale"])
            refresh(root, [sys.executable, PRODUCER, "second"])
            result = inspect(root)
            self.assertEqual(result["last_good"]["report"], "second")
            self.assertFalse(result["stale"])
            self.assertTrue(inspect(root, now=result["last_good"]["completed_at"] + 3601)["stale"])
            before = (root / "state.json").read_bytes()
            with (root / "writer.lock").open("a") as lock:
                fcntl.flock(lock, fcntl.LOCK_EX | fcntl.LOCK_NB)
                with self.assertRaises(OSError):
                    refresh(root, [sys.executable, PRODUCER, "blocked"])
            self.assertEqual((root / "state.json").read_bytes(), before)
            result["latest_attempt"]["status"] = "running"
            save(root, {key: result[key] for key in ("last_good", "latest_attempt")})
            before = (root / "state.json").read_bytes()
            self.assertTrue(inspect(root)["stale"])
            self.assertEqual((root / "state.json").read_bytes(), before)


if __name__ == "__main__":
    unittest.main()

Run the checks:

python3 -m unittest -v test_report_store

The test compares the entire saved good record after every failed refresh, including its timestamp. It also checks replacement, age expiry, lock contention, and an unfinished attempt. That last check simulates persisted running state; it is not a power-loss test.

Gotchas

A terminated runner can leave running behind. Read that as unconfirmed, not proof that a process is alive. The next authorized refresh can supersede it under the lock. If you need every historical attempt, add an append-only history through your existing storage mechanism.

Storage errors propagate. If writing the final snapshot fails, the last persisted view may still be running; it must not become a successful refresh in your scheduler’s monitoring. This demo does not implement disk-full recovery, schema migration, or a durable transaction spanning external files.

The producer is trusted and writes a small regular file. The size check bounds what the reader consumes, not the producer’s disk usage, descendants, or filesystem access. Subprocess timeouts here concern the direct child. A service with untrusted jobs needs its own process and storage isolation.

For an existing application, preserve its report format and reader API. If reports are large, store immutable artifacts and atomically replace a small pointer plus status snapshot, with a cleanup policy that cannot delete a referenced artifact. Start by adding the failed-retry test to the real refresh path.

Sources