Pular para o conteúdo principal

Retention

What ages out automatically, what does not, and how to plan disk capacity.

At a glance​

WhatAuto-pruned?How
Log files (/var/log/bulksigner/bulksigner-*.log)YesThe file sink rotates daily and retains 14 files (Logging:File:RetainedFileCountLimit).
data/input/ filesNoRemoved only after a successful sign-verify-promote, or by operator action.
data/processing/ directoriesNoCreated and removed by the worker per job. Lingering directories belong to failed/crashed jobs and are moved to error/ by the startup recovery sweep.
data/output/ files (signed or .enc envelopes, and a rejected file's .reject hand-back)NoThe cleanup action is currently a no-op stub. Files leave only by operator order — Clear Jobs, or deleting a job.
data/error/ directoriesNoSame.
Job / history / event rows in the operational storeNoSame — except by operator order: Clear Jobs deletes every job row and every operational event recorded before the clear started, leaving the JobsCleared event; deleting a job removes that job's rows and keeps every event.
Frozen approval rules and recorded approvalsNoNever pruned. Who authorised a payment, and under what rule, is exactly what an audit asks for after the fact, so they are retained once the job is terminal — until an operator clears or deletes the job, when they go with it.
Lacuna Signer documents a SignerFolder profile has taken (one row per document and flow action: the two ids at Signer, the job the discovery created, when)NoRetained for ever, deliberately, and not removed by deleting a job or by Clear Jobs either — the one record about a job that outlives it. It is what keeps a document this host already took from being signed a second time by the next sweep once its job is gone, so removing one would reopen that document. No personal data: two ids, a GUID and a timestamp. Empty on every deployment with no SignerFolder profile (2.16.0).
CNAB240 line detail (one row per payment)YesDeleted at the transition into Completed, Failed or Canceled. The one exception — see below.

What is retained is unchanged; where it is retained can differ​

Everything in that table is true on both database providers. Choosing Database:Provider = SqlServer moves the rows out of the SQLite file and into your own SQL Server or Azure SQL database; it changes nothing about which of them prune themselves, when, or why.

Two consequences follow from where, and both are yours rather than the service's:

  • Backup and size are your DBMS regime's concern under SqlServer — see Backup discipline below.
  • Switching database provider does not carry the record across. There is no importer. A deployment that switches comes up against an empty store — including the frozen approval rules and the recorded approvals, the two things this table keeps for ever precisely because they are the evidence of who authorised a payment file. Archive the old db/bulksigner.db deliberately, before the switch: Installation.

Logs — what the file sink does​

The file sink is configured under Logging:File:* (see Configuration):

KnobDefaultEffect
RollingIntervalDayA new file is created at the start of each UTC day.
FileSizeLimitBytes50 MBIf a file hits this size before the day rolls, the sink rolls to a sibling file.
RetainedFileCountLimit14Older rolled files are deleted by the sink.
MinimumLevelInformationAnything below this level is filtered out before reaching the file.

Net effect at the defaults: ~14 days of structured logs at ≤ 50 MB per day-file. Raise RetainedFileCountLimit for a longer forensic window or lower it for constrained disks. The sink flushes frequently, so concurrent readers (tail -f, journalctl -fu bulksigner) see writes in near-real-time.

Logs in a table — nothing prunes them​

Logging:AzureTable:* sends the same log events to an Azure Storage table so the diagnostic stream survives a host whose disk does not (see Configuration). It is the one destination in this product with no retention mechanism at all, and that is worth settling before you enable it rather than after.

Decide the pruning story before you turn the sink on

The file sink deletes its own old files (RetainedFileCountLimit). The table does not, and no Azure mechanism can do it for you: Azure Storage tables have no TTL, no lifecycle-management rule and no bulk-delete operation. The table grows for as long as the sink is enabled, and every row is billed storage plus the transactions to remove it later.

What that means in practice:

File sinkTable sink
Old data removed byThe sink itself, as it rollsNothing. You schedule a job, or it grows forever
Bounded byRetainedFileCountLimit × FileSizeLimitBytesYour own pruning schedule
Cost of leaving it aloneNil — self-limitingGrows monotonically

The supported answer is the Prune-BulkSignerLogTable.ps1 script in the deployment package, run on a schedule (an Azure Automation runbook, a scheduled task, or a container job) with a retention window that matches whatever your file sink keeps. Two operational notes:

  • Deletion is per entity. There is no DELETE WHERE, so pruning is a query followed by batched entity deletes, and the cost scales with what you are removing. Pruning weekly from the start is much cheaper than pruning once after a year.
  • Give each deployment its own table if two share a storage account. Distinguishing them by a column inside one table breaks the pruning recipe, which partitions by date rather than by deployment.

Cluster mode makes this sink close to mandatory — a Linux container's disk vanishes on recycle and takes rolled log files with it — which is why leaving it off there logs a Critical at startup rather than passing silently. That is also why the ordering matters: the deployment most likely to need the sink is the one least likely to have a pruning job yet. See High availability.

Operational data — not auto-pruned​

The cleanup action (POST /api/cleanup and the dashboard's System-page Cleanup button) is currently a no-op stub: it returns successfully with a "retention policy is not configured" message and removes nothing.

Why a stub and not "delete by age" by default?​

The shape of retention is deliberately operator-controlled. The audit trail (output/, error/, job history rows) is valuable for compliance: deleting a signed PDF a downstream verifier may still want to fetch, or a history row an auditor may still want to read, is a destructive action that should reflect a deliberate operator policy — not a default that surprises someone six months in.

Default behavior:

  • Signed outputs accumulate in output/. Operators or downstream automation move them out.
  • Error directories accumulate in error/. Operators inspect, then delete with normal filesystem commands.
  • Job rows accumulate in the operational store, which grows linearly with throughput — the SQLite file under Sqlite, your own database under SqlServer. There is no age-based pruning of job history in this version; the two operator actions below are the only things that remove it.

What an operator can delete: Clear Jobs and job deletion​

Nothing ages out, but two deliberate operator actions do delete — and both take files as well as rows.

Changed in 2.9.0 and 2.10.0 — Clear Jobs takes everything

Clear Jobs used to delete only finished job records and leave every file and every operational event in place. Since 2.9.0 it deletes every job record, whatever its status, and every file those jobs left behind; since 2.10.0 it deletes the operational events too.

  • Clear Jobs (System → Danger zone, or DELETE /api/jobs) deletes every job in every status — with its history, its frozen approval rule and recorded approvals, its CNAB240 line detail and its stage timings — and each job's files: the input, its processing/ and error/ folders, the signed output (or .enc envelope) and a rejected file's .reject hand-back. Every operational event recorded before the clear started goes with them, and one JobsCleared event is written as the record of the cut, naming who cleared and how many jobs, files, folders and events went. Pipeline state, signing profiles, configuration and log files are untouched. The response carries the counts (deleted, filesDeleted, foldersDeleted, eventsDeleted, itemsFailed); a file somebody else holds, or a folder the storage refuses, is left in place, counted and named in the log.
  • Deleting a job (the row's delete button on /jobs, one job at a time, behind a confirmation with an optional reason; there is no REST route) removes that job's rows — the same set as above — plus its processing/ and error/ folders, the output it recorded writing, and its input only if the job staged it and it is unchanged since. A job a worker is running cannot be deleted. Every operational event stays, and a JobDeleted event is added summarising what was removed, what was kept and any approvals the job carried.

Both are irreversible. Take a database backup first if the audit trail they remove still matters — the backup is the only copy of the trail that survives them.

The one exception: CNAB240 line detail​

The line-level parse of a payment file is the first and only operational data in the product that prunes itself. That is a deliberate departure from the stance above, and it is narrow on purpose.

When a job is parsed as a CNAB240 remessa, the pipeline stores one row per payment — carrying the beneficiary's name, their CPF/CNPJ where the file states one, and the destination account — in a 1:1 table beside the job. It serves two screens while the job is in flight and somebody may still act on it: the Payments table on /jobs/{id}, and the same table on the approval page.

The row is deleted at the transition into Completed, Failed or Canceled. Not on a schedule, not by the cleanup endpoint — at the transition itself, so there is no sweeper to fall behind and no window in which a terminal job still carries the data. Every route to a terminal status purges, including an operator's cancel, an approver's rejection, and an approval window expiring.

Two reasons, and the first is why this does not contradict the stance above:

  1. It is redundant once the job is terminal, not merely old. Everything else in the retention table is the only copy of what it records — delete a history row and the audit trail has a hole. The line detail is a cache of what is already in the file, and the file survives every terminal outcome: output/ when the job completes, output/ again — under a .reject name — when an approver vetoed it, and error/ when it fails, when a wait budget expires, or when an operator cancels it. The job also keeps its content SHA-256, so the surviving artifact can be proved to be the one that was parsed. Nothing becomes unknowable.
  2. It is the largest concentration of personal data the product holds — every beneficiary in every payroll, accumulating forever, with no remaining consumer once the job is done. An LGPD exposure that grows with throughput and buys nothing.

What is not touched by the purge: the job's summary figures (total, payment and cancellation counts, payment-date range), the content hash, and the job history. Those are permanent. The Payments panel says so plainly when the row is gone, rather than rendering an empty table that reads as data loss.

If a deployment needs the line detail to outlive the job, the artifact in output/ is the source of truth — archive that, not the database row.

Estimating disk growth​

Rough ballpark for a single instance:

Per-job artifactTypical size
Cleartext signed PDF~ source size + signature dictionary (~10–50 KB)
BSENC v1 envelopesource size + 37 bytes
Job row~ 1 KB
History row~ 200–500 bytes; 2–4 per successful job, more for retries / failures

For 10 000 jobs / day on average documents, expect roughly:

Surface30-day growth
output/dominated by document size (10 000 × 30 × source size)
db/bulksigner.db< 100 MB (rows are small)
logs/bounded by RetainedFileCountLimit × FileSizeLimitBytes (= 700 MB at defaults)
error/proportional to failure rate; usually small

The DB file rarely becomes the bottleneck. The output tree is the big surface — plan disk capacity (or external archival) accordingly.

Manual retention recipes​

Operators script their own retention. A few patterns:

# Linux: nightly cron that moves files older than 7 days into an archive tree.
find /var/lib/bulksigner/output -type f -mtime +7 \
-exec mv {} /archive/bulksigner/output/ \;
# Windows: scheduled task that moves files older than 7 days.
Get-ChildItem C:\ProgramData\Lacuna\BulkSigner\data\output `
-Recurse -File | Where-Object { $_.LastWriteTime -lt (Get-Date).AddDays(-7) } |
Move-Item -Destination D:\archive\bulksigner\output\

Moving (not deleting) preserves the audit trail in an off-instance location.

Prune error/ after triage​

# Delete error/ directories older than 30 days. Review first.
find /var/lib/bulksigner/error -mindepth 1 -maxdepth 1 -type d -mtime +30 -print
# review the output, then drop the -print and add -exec rm -rf {} \;

Prefer manual review here — error/ often contains the only forensic copy of what went wrong.

Trim history rows​

The integrity of the audit trail depends on the full history chain. If row volume becomes an operational problem, prefer archiving the SQLite file (mv bulksigner.db bulksigner-2026Q1.db, restart with a fresh DB) over partial deletes.

Backup discipline​

The built-in backup feature — SQLite only​

Under Database:Provider = Sqlite, the product can back the store up for you: Backup:Enabled = true adds a /backup dashboard page, GET|POST /api/backup, and an optional scheduler (Backup:IntervalHours). Artifacts go to a local path, an S3 or S3-compatible bucket, or an Azure Blob container, with Backup:RetainCount bounding how many are kept. Every key is in Configuration.

Backup:Enabled = true under SqlServer refuses the boot

It is a refusal naming both keys, not a silent no-op — because backing up a customer's own DBMS is that DBMS regime's job, and a feature that quietly did nothing would read as a backup that exists. Since cluster mode requires SqlServer, the combination is unreachable there by construction; Azure SQL's own point-in-time restore is the answer on that topology.

Two constraints on Backup:Disk:Path are worth repeating here because they are retention mistakes rather than configuration ones, and both are refused at boot: a path inside a watched input folder (the pipeline would ingest, sign and then delete your backup) and a path inside processing/, output/, error/ or db/.

Independent of retention, and of that feature​

  • Back up the operational store before every service upgrade. Schema migrations run automatically at startup and are one-way, on both providers — most releases since 2.0.0 add one. Take one before a Clear Jobs too: it is the only copy of the audit trail that survives it.
  • Snapshot output/ if it carries audit-significant artifacts. Especially when encryption is enabled — losing an encrypted file is doubly irrecoverable (no password = no plaintext).
  • Treat data/ as a unit when backing up. input/, processing/, output/, error/, db/, logs/ together describe the full operational state. A snapshot is consistent if taken with the service stopped or paused (and the in-flight job count at zero).
  • Under Database:Provider = SqlServer, the store is not in data/ and is your DBMS regime's concern — which is one of the two reasons a customer chooses that provider. Back it up on the same schedule as any other database of record, and keep the file-tree snapshot in step with it: a restored store whose processing/ directories no longer exist is a startup recovery sweep with nothing to reconcile against.
  • Under Storage:Provider = AzureFiles, processing/, output/ and error/ are not in data/ either. Back the share up through Azure Files' own snapshot or backup facilities; logs/ and, under Sqlite, db/ remain on the host.

See Operations for the pause / upgrade / backup procedure.


Next: Troubleshooting. Previous: Approvals.