Retention
What ages out automatically, what does not, and how to plan disk capacity.
At a glance
| What | Auto-pruned? | How |
|---|---|---|
Log files (/var/log/bulksigner/bulksigner-*.log) | Yes | The file sink rotates daily and retains 14 files (Logging:File:RetainedFileCountLimit). |
data/input/ files | No | Removed only after a successful sign-verify-promote, or by operator action. |
data/processing/ directories | No | Created and removed by the worker per job. Lingering directories belong to failed/crashed jobs and are moved to error/ by the startup recovery sweep. |
data/output/ files (signed or .enc envelopes, and a rejected file's .reject hand-back) | No | The cleanup action is currently a no-op stub. Files leave only by operator order — Clear Jobs, or deleting a job. |
data/error/ directories | No | Same. |
| Job / history / event rows in the operational store | No | Same — except by operator order: Clear Jobs deletes every job row and every operational event recorded before the clear started, leaving the JobsCleared event; deleting a job removes that job's rows and keeps every event. |
| Frozen approval rules and recorded approvals | No | Never pruned. Who authorised a payment, and under what rule, is exactly what an audit asks for after the fact, so they are retained once the job is terminal — until an operator clears or deletes the job, when they go with it. |
Lacuna Signer documents a SignerFolder profile has taken (one row per document and flow action: the two ids at Signer, the job the discovery created, when) | No | Retained for ever, deliberately, and not removed by deleting a job or by Clear Jobs either — the one record about a job that outlives it. It is what keeps a document this host already took from being signed a second time by the next sweep once its job is gone, so removing one would reopen that document. No personal data: two ids, a GUID and a timestamp. Empty on every deployment with no SignerFolder profile (2.16.0). |
| CNAB240 line detail (one row per payment) | Yes | Deleted at the transition into Completed, Failed or Canceled. The one exception — see below. |
What is retained is unchanged; where it is retained can differ
Everything in that table is true on both database providers. Choosing
Database:Provider = SqlServer moves the rows out of
the SQLite file and into your own SQL Server or Azure SQL database; it changes nothing about which of
them prune themselves, when, or why.
Two consequences follow from where, and both are yours rather than the service's:
- Backup and size are your DBMS regime's concern under
SqlServer— see Backup discipline below. - Switching database provider does not carry the record across. There is no importer. A deployment
that switches comes up against an empty store — including the frozen approval rules and the
recorded approvals, the two things this table keeps for ever precisely because they are the evidence
of who authorised a payment file. Archive the old
db/bulksigner.dbdeliberately, before the switch: Installation.
Logs — what the file sink does
The file sink is configured under Logging:File:* (see
Configuration):
| Knob | Default | Effect |
|---|---|---|
RollingInterval | Day | A new file is created at the start of each UTC day. |
FileSizeLimitBytes | 50 MB | If a file hits this size before the day rolls, the sink rolls to a sibling file. |
RetainedFileCountLimit | 14 | Older rolled files are deleted by the sink. |
MinimumLevel | Information | Anything below this level is filtered out before reaching the file. |
Net effect at the defaults: ~14 days of structured logs at ≤ 50 MB per day-file. Raise
RetainedFileCountLimit for a longer forensic window or lower it for constrained disks. The sink
flushes frequently, so concurrent readers (tail -f, journalctl -fu bulksigner) see writes in
near-real-time.
Logs in a table — nothing prunes them
Logging:AzureTable:* sends the same log events to an Azure Storage table so the diagnostic stream
survives a host whose disk does not (see
Configuration). It is the one destination in
this product with no retention mechanism at all, and that is worth settling before you enable it
rather than after.
The file sink deletes its own old files (RetainedFileCountLimit). The table does not, and no Azure
mechanism can do it for you: Azure Storage tables have no TTL, no lifecycle-management rule and no
bulk-delete operation. The table grows for as long as the sink is enabled, and every row is billed
storage plus the transactions to remove it later.
What that means in practice:
| File sink | Table sink | |
|---|---|---|
| Old data removed by | The sink itself, as it rolls | Nothing. You schedule a job, or it grows forever |
| Bounded by | RetainedFileCountLimit × FileSizeLimitBytes | Your own pruning schedule |
| Cost of leaving it alone | Nil — self-limiting | Grows monotonically |
The supported answer is the Prune-BulkSignerLogTable.ps1 script in the deployment package, run on a
schedule (an Azure Automation runbook, a scheduled task, or a container job) with a retention window
that matches whatever your file sink keeps. Two operational notes:
- Deletion is per entity. There is no
DELETE WHERE, so pruning is a query followed by batched entity deletes, and the cost scales with what you are removing. Pruning weekly from the start is much cheaper than pruning once after a year. - Give each deployment its own table if two share a storage account. Distinguishing them by a column inside one table breaks the pruning recipe, which partitions by date rather than by deployment.
Cluster mode makes this sink close to mandatory — a Linux container's disk vanishes on recycle and takes rolled log files with it — which is why leaving it off there logs a Critical at startup rather than passing silently. That is also why the ordering matters: the deployment most likely to need the sink is the one least likely to have a pruning job yet. See High availability.
Operational data — not auto-pruned
The cleanup action (POST /api/cleanup and the dashboard's System-page Cleanup button) is currently
a no-op stub: it returns successfully with a "retention policy is not configured" message and
removes nothing.
Why a stub and not "delete by age" by default?
The shape of retention is deliberately operator-controlled. The audit trail (output/, error/, job
history rows) is valuable for compliance: deleting a signed PDF a downstream verifier may still
want to fetch, or a history row an auditor may still want to read, is a destructive action that
should reflect a deliberate operator policy — not a default that surprises someone six months in.
Default behavior:
- Signed outputs accumulate in
output/. Operators or downstream automation move them out. - Error directories accumulate in
error/. Operators inspect, then delete with normal filesystem commands. - Job rows accumulate in the operational store, which grows linearly with throughput — the SQLite file
under
Sqlite, your own database underSqlServer. There is no age-based pruning of job history in this version; the two operator actions below are the only things that remove it.
What an operator can delete: Clear Jobs and job deletion
Nothing ages out, but two deliberate operator actions do delete — and both take files as well as rows.
Clear Jobs used to delete only finished job records and leave every file and every operational event in place. Since 2.9.0 it deletes every job record, whatever its status, and every file those jobs left behind; since 2.10.0 it deletes the operational events too.
- Clear Jobs (System → Danger zone, or
DELETE /api/jobs) deletes every job in every status — with its history, its frozen approval rule and recorded approvals, its CNAB240 line detail and its stage timings — and each job's files: the input, itsprocessing/anderror/folders, the signed output (or.encenvelope) and a rejected file's.rejecthand-back. Every operational event recorded before the clear started goes with them, and oneJobsClearedevent is written as the record of the cut, naming who cleared and how many jobs, files, folders and events went. Pipeline state, signing profiles, configuration and log files are untouched. The response carries the counts (deleted,filesDeleted,foldersDeleted,eventsDeleted,itemsFailed); a file somebody else holds, or a folder the storage refuses, is left in place, counted and named in the log. - Deleting a job (the row's delete button on
/jobs, one job at a time, behind a confirmation with an optional reason; there is no REST route) removes that job's rows — the same set as above — plus itsprocessing/anderror/folders, the output it recorded writing, and its input only if the job staged it and it is unchanged since. A job a worker is running cannot be deleted. Every operational event stays, and aJobDeletedevent is added summarising what was removed, what was kept and any approvals the job carried.
Both are irreversible. Take a database backup first if the audit trail they remove still matters — the backup is the only copy of the trail that survives them.
The one exception: CNAB240 line detail
The line-level parse of a payment file is the first and only operational data in the product that prunes itself. That is a deliberate departure from the stance above, and it is narrow on purpose.
When a job is parsed as a CNAB240 remessa, the pipeline stores one row per payment —
carrying the beneficiary's name, their CPF/CNPJ where the file states one, and the destination account
— in a 1:1 table beside the job. It serves two screens while the job is in flight and somebody may
still act on it: the Payments table on /jobs/{id}, and the same table on the
approval page.
The row is deleted at the transition into Completed, Failed or Canceled. Not on a schedule,
not by the cleanup endpoint — at the transition itself, so there is no sweeper to fall behind and no
window in which a terminal job still carries the data. Every route to a terminal status purges,
including an operator's cancel, an approver's rejection, and an approval window expiring.
Two reasons, and the first is why this does not contradict the stance above:
- It is redundant once the job is terminal, not merely old. Everything else in the retention table
is the only copy of what it records — delete a history row and the audit trail has a hole. The
line detail is a cache of what is already in the file, and the file survives every terminal outcome:
output/when the job completes,output/again — under a.rejectname — when an approver vetoed it, anderror/when it fails, when a wait budget expires, or when an operator cancels it. The job also keeps its content SHA-256, so the surviving artifact can be proved to be the one that was parsed. Nothing becomes unknowable. - It is the largest concentration of personal data the product holds — every beneficiary in every payroll, accumulating forever, with no remaining consumer once the job is done. An LGPD exposure that grows with throughput and buys nothing.
What is not touched by the purge: the job's summary figures (total, payment and cancellation counts, payment-date range), the content hash, and the job history. Those are permanent. The Payments panel says so plainly when the row is gone, rather than rendering an empty table that reads as data loss.
If a deployment needs the line detail to outlive the job, the artifact in output/ is the source of
truth — archive that, not the database row.
Estimating disk growth
Rough ballpark for a single instance:
| Per-job artifact | Typical size |
|---|---|
| Cleartext signed PDF | ~ source size + signature dictionary (~10–50 KB) |
| BSENC v1 envelope | source size + 37 bytes |
| Job row | ~ 1 KB |
| History row | ~ 200–500 bytes; 2–4 per successful job, more for retries / failures |
For 10 000 jobs / day on average documents, expect roughly:
| Surface | 30-day growth |
|---|---|
output/ | dominated by document size (10 000 × 30 × source size) |
db/bulksigner.db | < 100 MB (rows are small) |
logs/ | bounded by RetainedFileCountLimit × FileSizeLimitBytes (= 700 MB at defaults) |
error/ | proportional to failure rate; usually small |
The DB file rarely becomes the bottleneck. The output tree is the big surface — plan disk capacity (or external archival) accordingly.
Manual retention recipes
Operators script their own retention. A few patterns:
Move-and-archive output/ (recommended)
# Linux: nightly cron that moves files older than 7 days into an archive tree.
find /var/lib/bulksigner/output -type f -mtime +7 \
-exec mv {} /archive/bulksigner/output/ \;
# Windows: scheduled task that moves files older than 7 days.
Get-ChildItem C:\ProgramData\Lacuna\BulkSigner\data\output `
-Recurse -File | Where-Object { $_.LastWriteTime -lt (Get-Date).AddDays(-7) } |
Move-Item -Destination D:\archive\bulksigner\output\
Moving (not deleting) preserves the audit trail in an off-instance location.
Prune error/ after triage
# Delete error/ directories older than 30 days. Review first.
find /var/lib/bulksigner/error -mindepth 1 -maxdepth 1 -type d -mtime +30 -print
# review the output, then drop the -print and add -exec rm -rf {} \;
Prefer manual review here — error/ often contains the only forensic copy of what went wrong.
Trim history rows
The integrity of the audit trail depends on the full history chain. If row volume becomes an
operational problem, prefer archiving the SQLite file (mv bulksigner.db bulksigner-2026Q1.db,
restart with a fresh DB) over partial deletes.
Backup discipline
The built-in backup feature — SQLite only
Under Database:Provider = Sqlite, the product can back the store up for you: Backup:Enabled = true
adds a /backup dashboard page, GET|POST /api/backup, and an optional scheduler
(Backup:IntervalHours). Artifacts go to a local path, an S3 or S3-compatible bucket, or an Azure Blob
container, with Backup:RetainCount bounding how many are kept. Every key is in
Configuration.
Backup:Enabled = true under SqlServer refuses the bootIt is a refusal naming both keys, not a silent no-op — because backing up a customer's own DBMS is that
DBMS regime's job, and a feature that quietly did nothing would read as a backup that exists. Since
cluster mode requires SqlServer, the combination is unreachable there by construction; Azure SQL's
own point-in-time restore is the answer on that topology.
Two constraints on Backup:Disk:Path are worth repeating here because they are retention mistakes
rather than configuration ones, and both are refused at boot: a path inside a watched input folder
(the pipeline would ingest, sign and then delete your backup) and a path inside processing/,
output/, error/ or db/.
Independent of retention, and of that feature
- Back up the operational store before every service upgrade. Schema migrations run automatically at startup and are one-way, on both providers — most releases since 2.0.0 add one. Take one before a Clear Jobs too: it is the only copy of the audit trail that survives it.
- Snapshot
output/if it carries audit-significant artifacts. Especially when encryption is enabled — losing an encrypted file is doubly irrecoverable (no password = no plaintext). - Treat
data/as a unit when backing up.input/,processing/,output/,error/,db/,logs/together describe the full operational state. A snapshot is consistent if taken with the service stopped or paused (and the in-flight job count at zero). - Under
Database:Provider = SqlServer, the store is not indata/and is your DBMS regime's concern — which is one of the two reasons a customer chooses that provider. Back it up on the same schedule as any other database of record, and keep the file-tree snapshot in step with it: a restored store whoseprocessing/directories no longer exist is a startup recovery sweep with nothing to reconcile against. - Under
Storage:Provider = AzureFiles,processing/,output/anderror/are not indata/either. Back the share up through Azure Files' own snapshot or backup facilities;logs/and, underSqlite,db/remain on the host.
See Operations for the pause / upgrade / backup procedure.
Next: Troubleshooting. Previous: Approvals.