A line number is not an identity
Shipped
My running coach keeps its saved preferences in a plain markdown file, one bullet per note, because I want to be able to open it in a text editor and fix a line by hand. Two of the tools that edit it, update_user_note and delete_user_note, used to address a note by its line index. This release replaces the index with an 8-character content handle, moves every write inside one held lock, and exits each write through a temp file and an atomic replace. The same release fixes a trailing-newline bug that could weld two notes into one, and makes recency follow the note’s timestamp rather than its position in the file.
The pattern underneath is not specific to fitness notes. Any time a program (or an agent) edits records in a file that a human also edits, three questions decide whether it is safe: what identifies a record, when the target is resolved relative to the write, and whether a reader can ever see the file half-written. This guide builds a small version of the answer.
The file, and why a line index goes wrong
The file looks like this. One record per line, a timestamp, a separator, the text:
- 2026-08-03T12:19:00 | Render distances in miles, never km
- 2026-08-10T07:41:12 | Roast me when I skip a long run
- 2026-08-24T18:02:55 | No emoji in the morning brief
Addressing a record as “line 2” works right up until something else moves. In the real file, a delete shifted every later index down by one, and a size-based rotation that archives the oldest notes renumbered the whole thing. Neither write path noticed. A caller that read line 2, decided to update it, and wrote line 2 could land on a different preference than the one it had read, and the tool reported success either way. The problem is not the file format. It is that the address was recomputed from position on every read, so it described the file at read time and nothing else.
A content handle fixes this by construction. Derive the address from the record’s own fields, so the only thing that can change it is the record itself changing, which is exactly the case where a stale caller should be refused.
Step 1: derive the handle from content
Create prefs.py. The handle is the first eight hex characters of a SHA-256 over the timestamp and the text. It is short on purpose. The only adversary is a stale value, and an agent has to transcribe it verbatim from a prompt.
# prefs.py
import fcntl
import hashlib
import os
import stat
import tempfile
from contextlib import contextmanager
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
@dataclass(frozen=True)
class Note:
handle: str
timestamp: str
text: str
def handle_for(timestamp: str, text: str) -> str:
return hashlib.sha256(f"{timestamp}\n{text}".encode()).hexdigest()[:8]
def parse_line(line: str) -> Note | None:
body = line.rstrip("\n")
if not body.startswith("- ") or " | " not in body:
return None
ts, text = body[2:].split(" | ", 1)
return Note(handle_for(ts, text), ts, text)
def read_notes(path: Path) -> list[Note]:
if not path.exists():
return []
lines = path.read_text(encoding="utf-8").splitlines(keepends=True)
return [n for ln in lines if (n := parse_line(ln)) is not None]
Two identical records (same timestamp, same text) collide. That is tolerable here: the writer acts on the first match and reports how many matched, and after one edit the handles are unique again. Do not reach for a stored id column to avoid it; the whole point is that a hand edit needs no bookkeeping.
Step 2: one lock, one read, one atomic write
The second question is when the target gets resolved. If the tool reads the file, resolves the handle, and then takes a lock to write, another writer can land between the read and the lock. So the lock goes first, the read happens inside it, resolution happens against that read, and the rewrite is assembled from the same lines.
Lock a sidecar file, not the file you are about to replace. The replace swaps the inode out from under any lock held on the old one, so a lock on the notes file itself protects the version that is about to disappear.
def current_umask() -> int:
m = os.umask(0)
os.umask(m)
return m
def write_atomic(path: Path, text: str, mode: int) -> None:
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix=path.name + ".")
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
f.write(text)
f.flush()
os.fsync(f.fileno())
os.chmod(tmp, mode)
os.replace(tmp, path)
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
class Rewrite:
def __init__(self, existing: str):
if existing and not existing.endswith("\n"):
existing += "\n"
self.lines: list[str] = existing.splitlines(keepends=True)
self.text: str | None = None
def resolve(self, handle: str) -> int | None:
for i, ln in enumerate(self.lines):
note = parse_line(ln)
if note is not None and note.handle == handle:
return i
return None
@contextmanager
def locked_rewrite(path: Path):
path.parent.mkdir(parents=True, exist_ok=True)
with open(path.with_name(path.name + ".lock"), "a+") as lock:
fcntl.flock(lock.fileno(), fcntl.LOCK_EX)
existing = path.read_text(encoding="utf-8") if path.exists() else ""
mode = (
stat.S_IMODE(path.stat().st_mode)
if path.exists()
else 0o666 & ~current_umask()
)
ctx = Rewrite(existing)
yield ctx
if ctx.text is not None:
if ctx.text and not ctx.text.endswith("\n"):
ctx.text += "\n"
write_atomic(path, ctx.text, mode)
Three details carry the weight. fcntl.flock with LOCK_EX is the exclusive lock, per the Python fcntl docs, and the with block releases it by closing the handle. The temp file lives in the same directory so the final os.replace is a rename on one filesystem; the rename(2) man page is explicit that when the destination exists it is replaced atomically and no other process ever finds it missing, and equally explicit that a rename across mount points fails with EXDEV. And the write follows the sequence Jeff Moyer lays out in Ensuring data reaches disk: temp file on the same filesystem, write, fsync, rename. His article also fsyncs the containing directory for durability across a crash; I left that out because these are preferences, not ledger rows, and you should add it if a lost write would matter.
The Rewrite context normalises a trailing newline on the way in and locked_rewrite normalises it again on the way out. That is not tidiness. It is the fix for a real bug I describe under Gotchas.
Step 3: writers that resolve inside the lock
Every writer now has the same shape: enter the lock, resolve against the lines that were just read, assign the new whole-file text, leave. Leaving ctx.text as None writes nothing, which is how a stale handle is refused without touching the file.
def now() -> str:
return datetime.now().replace(microsecond=0).isoformat()
def append_note(path: Path, text: str) -> Note:
ts = now()
with locked_rewrite(path) as ctx:
ctx.text = "".join(ctx.lines) + f"- {ts} | {text}\n"
return Note(handle_for(ts, text), ts, text)
def update_note(path: Path, handle: str, text: str) -> Note | None:
ts = now()
with locked_rewrite(path) as ctx:
i = ctx.resolve(handle.strip().strip("[]").lower())
if i is None:
return None
ctx.lines[i] = f"- {ts} | {text}\n"
ctx.text = "".join(ctx.lines)
return Note(handle_for(ts, text), ts, text)
def delete_note(path: Path, handle: str) -> bool:
with locked_rewrite(path) as ctx:
i = ctx.resolve(handle.strip().strip("[]").lower())
if i is None:
return False
del ctx.lines[i]
ctx.text = "".join(ctx.lines)
return True
The handle normalisation (strip, drop brackets, lowercase) exists because the agent reads [a1b2c3d4] off a rendered prompt and often echoes the brackets back. Matching is still exact; there is no prefix matching, because a prefix match is a second way to land on the wrong record.
Use it, then check the two properties
Save this as demo.py next to prefs.py:
# demo.py
from pathlib import Path
from prefs import append_note, delete_note, read_notes, update_note
path = Path("notes.md")
path.unlink(missing_ok=True)
a = append_note(path, "Render distances in miles, never km")
b = append_note(path, "Roast me when I skip a long run")
c = append_note(path, "No emoji in the morning brief")
print("delete a:", delete_note(path, a.handle))
print("update b:", update_note(path, b.handle, "Roast me when I skip any planned run") is not None)
print("stale b: ", update_note(path, b.handle, "must not land"))
for n in read_notes(path):
print(n.handle, n.timestamp, n.text)
Run python demo.py. The timestamps and handles will differ from mine, but the shape should be this: the delete succeeds, the first update succeeds, the second update with the now-stale handle returns None and writes nothing, and two notes remain.
delete a: True
update b: True
stale b: None
24d11554 2026-09-05T20:30:50 Roast me when I skip any planned run
5c549ab0 2026-09-05T20:30:50 No emoji in the morning brief
Before the fix, that third call would have re-resolved “the note at position 0”, which after the delete was a different preference, and overwritten it.
The second property is that a reader never observes a half-written file. Save this as race.py:
# race.py
import threading
from pathlib import Path
from prefs import append_note, read_notes
path = Path("race.md")
path.unlink(missing_ok=True)
append_note(path, "seed")
empty = 0
def reader() -> None:
global empty
for _ in range(300):
if not read_notes(path):
empty += 1
def writer() -> None:
for i in range(300):
append_note(path, f"note {i}")
threads = [threading.Thread(target=reader), threading.Thread(target=writer)]
for t in threads:
t.start()
for t in threads:
t.join()
print("empty reads:", empty, "of 300")
You should see empty reads: 0 of 300. Swap write_atomic for a plain path.write_text(text) and run it a few times; a truncate-then-write shows up as non-zero. In the real code, before this release, that window was measured at 22% of reads in a two-thread loop, and the reader on the other side was the coach-persona build, which quietly produced a prompt with no saved-preferences section at all.
Gotchas
The temp file has the wrong permissions. tempfile.mkstemp creates a file readable and writable only by the creating user, per the tempfile docs. Rename that over a file that was 0644 and the file is now 0600; the next reader running as a different user gets a permission error, and nothing in your code looks wrong. Read the original mode inside the lock and chmod the temp file to it before the replace. That is what the mode argument is for.
A missing trailing newline welds two records into one. A file that a human edits will eventually be saved without its final newline. The old rotation code joined the surviving lines and appended the new one; the old archive-append checked the newline of the text it was writing but not of the file it was appending to. Either gap merged the newest surviving note into the incoming one, unarchived and unrecoverable. The escape is to establish the boundary once, in the rewrite context, so no writer has to remember it.
Recency was file position, and the rotation trusted it. Updating a note in place refreshes its timestamp, so file order stopped being recency order the moment any note was refined. The 4 KB rotation that evicts the oldest notes evicted by position, which means it archived the freshest note first. Rank by parsed timestamp, treat an absent or malformed timestamp as oldest, and make sure the rotation triggered by an update can never evict the record that update just rewrote.
Locking the file you are replacing. The obvious lock target is the notes file. But os.replace swaps a new inode into that name, so the lock is now on an orphan. Lock a sidecar (notes.md.lock) that nothing ever replaces.
Sources
- rename(2) man page — atomic replacement when the destination exists;
EXDEVacross mount points - Python fcntl module —
flockand theLOCK_EX/LOCK_SH/LOCK_NBoperations - Ensuring data reaches disk, Jeff Moyer, LWN — the temp file, fsync, rename sequence and the directory fsync
- Python tempfile module —
mkstempsemantics and the creating-user-only permissions
Changelog
- update_user_note / delete_user_note can silently hit the wrong preference — _locked_rewrite — read, resolve and rewrite inside one held sidecar… (#226) (a692563)
- notes: line framing + content-handle addressing (#132, layers 2–3) (#233) (393ec40)
- notes: recency ordering + cap enforcement on the update path (#132, layers 4–5) (#234) (70a74ac)
- observations: validate obs_type and observation_id; 0.61.0 (#132, layer 6) (#235) (7326f66)