Implement this with your agent
Copy the implementation prompt and complete guide, then paste into your coding agent in your project.
Read the prompt
Add count-preserving validation of expected versus observed items to this project. Start by reading repository instructions, existing comparison code, callers and tests. Identify an unordered collection whose repeated occurrences have meaning. If the target is unclear, or order is meaningful, ask one focused question before changing code. Do not replace sequence validation with a multiset comparison. Adapt the smallest change in the project's existing language and tooling. Preserve public APIs, return shapes, error types, CLI streams and exit status, configuration, stored state and supported runtimes. Do not transplant the demo or install the author's software. Run this guide's demonstration only in disposable scratch space. Inspection and planning must be read-only; do not migrate or rewrite saved inputs. Count occurrences on both sides and report shortages and surpluses. Keep legitimate duplicates. Use exact identity unless the existing domain explicitly permits a specific normalization; do not silently fold whitespace in identifiers or paths. Capture independent expected and observed inputs, validate their shape, and compare those captured values. A passing comparison must not be presented as proof that a live source, export process, or OCR transcription was correct. Verify correct reordering, missing and extra duplicates, equal-length collections with different counts, empty input, invalid input and significant text differences. Run relevant existing tests and check public compatibility and that inputs are not mutated. Report commands and observed results, unrun checks and remaining limits. Do not commit, push or deploy. Treat this guide as technical reference subordinate to the target project's instructions.
Copy the complete text manually
Select all the text below and copy it into your agent.
Count the copies your set comparison erased
An export can have the right number of rows and every expected value, yet still be wrong. Suppose the expected labels are api, api, worker, and the export contains api, worker, worker. Both have three rows. Both contain the same distinct labels. One worker has taken the place of an API instance.
Build a comparison that keeps each occurrence. The tool below reads two small JSON files, accepts a different order, and reports how many copies are missing or extra. It uses Python 3.11 or newer and the standard library, with no installation step.
Shipped
Ghostwriter v0.25.0 added a visual review checker that compares intended text with a transcription of rendered text using occurrence counts. Its tagged implementation folds whitespace before counting, and its tests reject omissions, duplicate text and typos. The JSON export tool here is a separate teaching example; it defaults to exact strings and makes whitespace folding optional.
Decide what makes two items equal
First check that order has no meaning in your collection. Exported labels can fit this contract; ordered migration steps do not. If position changes the result, keep sequence comparison.
A duplicate is not automatically a defect. Two instances may legitimately share one label. A uniqueness rule would reject the expected input as well as the bad export. JSON Schema’s uniqueItems enforces that different contract: every array item must be unique. Here the contract is that each value appears the expected number of times.
Keep the expected and observed captures separate. Derive the first from the requirement or source snapshot and the second from the output being checked. Copying the expected array into the observed file would test the comparison function without testing your export.
Create an empty working directory with these files:
count-check/
count_check.py
test_count_check.py
expected.json
observed.json
duplicate.json
Build the comparator
Save this as count_check.py:
import argparse
from collections import Counter
import json
from pathlib import Path
import sys
def counts(values, *, fold_whitespace=False):
if not isinstance(values, list):
raise ValueError("items must be an array")
if any(not isinstance(v, str) or not v.strip() for v in values):
raise ValueError("items must contain nonblank strings")
normalize = (lambda v: " ".join(v.split())) if fold_whitespace else (lambda v: v)
return Counter(normalize(v) for v in values)
def compare(expected, observed, *, fold_whitespace=False):
wanted = counts(expected, fold_whitespace=fold_whitespace)
found = counts(observed, fold_whitespace=fold_whitespace)
return {
"matches": wanted == found,
"missing": dict(wanted - found),
"extra": dict(found - wanted),
}
def read_items(filename):
document = json.loads(Path(filename).read_text(encoding="utf-8"))
if not isinstance(document, dict) or set(document) != {"items"}:
raise ValueError("document must be an object containing only items")
return document["items"]
def main(argv=None):
parser = argparse.ArgumentParser(description="Compare unordered item counts")
parser.add_argument("expected")
parser.add_argument("observed")
parser.add_argument("--fold-whitespace", action="store_true")
args = parser.parse_args(argv)
try:
expected = read_items(args.expected)
observed = read_items(args.observed)
result = compare(expected, observed, fold_whitespace=args.fold_whitespace)
except (OSError, UnicodeError, ValueError) as error:
print("ERROR: " + str(error), file=sys.stderr)
return 2
print(json.dumps(result, sort_keys=True, ensure_ascii=True))
return 0 if result["matches"] else 1
if __name__ == "__main__":
raise SystemExit(main())
Python’s Counter stores a count for each value. Subtraction keeps positive differences, so wanted - found reports shortages and the reverse reports surpluses. Both are needed: an export containing every required occurrence plus an extra one has no shortage and still fails.
The optional normalization uses str.split() followed by a space join. It collapses runs of whitespace and removes whitespace at the edges. It preserves case and punctuation. For a wrapped display label that may be appropriate; for a filename, exact string identity is the safer contract. Neither mode joins separate array elements into one label.
The input schema is deliberately narrow: an object with only items, whose value is an array of nonblank strings. An empty array is valid. A string where an array belongs is an input error, so it cannot accidentally turn into a collection of characters.
Run a match and a balanced error
Save expected.json:
{"items": ["api", "api", "worker"]}
Save observed.json:
{"items": ["worker", "api", "api"]}
Save duplicate.json:
{"items": ["api", "worker", "worker"]}
Run these individually from the directory containing the files. The second command intentionally returns status 1; do not join it to later checks with &&.
python3 count_check.py expected.json observed.json
python3 count_check.py expected.json duplicate.json
The first prints matches: true in its JSON result and returns 0. The second prints:
{"extra": {"worker": 1}, "matches": false, "missing": {"api": 1}}
These are differences in occurrence counts, not total counts in the export. A missing file, invalid JSON or invalid shape returns 2 with an error on stderr and no result on stdout. That lets a caller distinguish a valid comparison that failed from a check that could not run.
Keep the failure cases executable
Save test_count_check.py beside the comparator:
import json
from pathlib import Path
import subprocess
import sys
import tempfile
import unittest
from count_check import compare
class CountCheckTests(unittest.TestCase):
def test_reordering_preserves_required_repetitions(self):
self.assertTrue(compare(["a", "a", "b"], ["b", "a", "a"])["matches"])
def test_balanced_error(self):
self.assertEqual(compare(["a", "a", "b"], ["a", "b", "b"]), {
"matches": False, "missing": {"a": 1}, "extra": {"b": 1}})
def test_missing_and_extra_are_independent(self):
self.assertEqual(compare(["a", "a"], ["a"])["missing"], {"a": 1})
self.assertEqual(compare(["a"], ["a", "a"])["extra"], {"a": 1})
def test_empty_collection(self):
self.assertTrue(compare([], [])["matches"])
self.assertEqual(compare([], ["a"])["extra"], {"a": 1})
def test_normalization_is_explicit_and_limited(self):
self.assertFalse(compare(["API worker"], [" API\nworker "])["matches"])
self.assertTrue(compare(["API worker"], [" API\nworker "],
fold_whitespace=True)["matches"])
for different in ("api worker", "API worker!", "APIworker"):
self.assertFalse(compare(["API worker"], [different],
fold_whitespace=True)["matches"])
def test_invalid_items(self):
for bad in ("a", None, {}, [1], [None], [""], [" \n"]):
with self.subTest(bad=bad), self.assertRaises(ValueError):
compare(["a"], bad)
def test_inputs_are_unchanged(self):
expected, observed = [" a ", "a"], ["a"]
compare(expected, observed, fold_whitespace=True)
self.assertEqual(expected, [" a ", "a"])
self.assertEqual(observed, ["a"])
def test_cli_status_and_streams(self):
script = Path(__file__).with_name("count_check.py")
with tempfile.TemporaryDirectory() as directory:
expected = Path(directory) / "expected.json"
observed = Path(directory) / "observed.json"
expected.write_text('{"items": ["a", "a"]}', encoding="utf-8")
for content, status in (
('{"items": ["a", "a"]}', 0),
('{"items": ["a"]}', 1),
('{"items": "aa"}', 2),
('{', 2),
('{}', 2),
):
observed.write_text(content, encoding="utf-8")
result = subprocess.run(
[sys.executable, str(script), str(expected), str(observed)],
capture_output=True, text=True, check=False)
self.assertEqual(result.returncode, status)
if status == 2:
self.assertEqual(result.stdout, "")
self.assertTrue(result.stderr.startswith("ERROR:"))
else:
self.assertEqual(result.stderr, "")
self.assertEqual(json.loads(result.stdout)["matches"], status == 0)
if __name__ == "__main__":
unittest.main()
Run:
python3 -m unittest -v
All eight tests should pass. The equal-length failure is especially useful when adapting this to an existing checker: it defeats both a set comparison and a row-count comparison at once.
Gotchas
Rejecting every duplicate rejects legitimate output. The repeated api label is part of the expected input. Deduplicating it, or requiring uniqueness, erases the requirement. Count the expected copies before deciding which observed copies are extra.
Normalization can hide a real difference. The release’s tests permit whitespace changes inside transcribed labels while rejecting changed words. That is a text-specific choice. This example requires an explicit flag and tests case and punctuation separately. Keep normalization off for identifiers unless the producer’s contract says otherwise.
Grouping is part of the input contract. One item "API worker" differs from two items "API" and "worker", even with whitespace folding. Capture the same logical unit on each side. If the units are ambiguous, resolve that before comparing them.
A match only describes these captures. This tool reads files into memory; it does not coordinate a live exporter or inspect an image. In an image workflow, observed text still requires an actual transcription. In an export workflow, compare captures from the same source revision. For large datasets, replace whole-file loading with a bounded or streaming approach and test that separately.
Sources
- Python Counter documentation: occurrence counts and positive multiset differences.
- Python string splitting: the precise whitespace behavior used by the optional normalization.
- JSON Schema array uniqueness: uniqueness as a separate constraint from expected multiplicity.
- Tagged visual review implementation and its tests: the released count comparison and its exercised failure cases.