The cache is where a declared resource becomes a local path. This article covers the layout, what makes a cached copy trustworthy, how two sessions avoid downloading the same file twice, and the retention policy CRAN requires of a package that writes to a user directory.
None of it is private. The cache is an ordinary directory tree, and that is a deliberate property: copying it to an offline machine, restoring it from a CI cache action, or looking at it with a file browser all work.
Where it lives
getaca_cache_dir()
#> [1] "<tempdir>/getaca-cache-vignette"That is the sandbox this vignette runs in. The default is
tools::R_user_dir("getaca", "cache"), the location CRAN
permits for cached downloads. Two overrides take precedence, in this
order:
options(getaca.cache = "/mnt/fast/getaca") # this session
Sys.setenv(GETACA_CACHE = "/mnt/fast/getaca") # this process and its childrenThe environment variable is the one to reach for in CI and in job scripts, because it survives into the R processes a build step spawns. Asking for the directory does not create it; the first successful retrieval does.
Layout
<cache>/
blobs/sha256/<aa>/<sha256> verified bytes, named by their own checksum
.locks/ one per checksum, held during a transfer
.tmp/ in-flight downloads, never visible as cache
<package>/
index.rds provenance for this package only
<name>/<version>/
raw/<file> this slot's name for a blob
proc-<processor-id>/ processed result, own provenance
Everything a package declares is scoped by declaring package, then
resource name, then version. That falls out of identity being the triple
package / name / version, and it buys two things. Two
packages declaring a resource called "backbone" never share
a slot, so one package’s registry update cannot affect another’s cached
data. And a version can never be overwritten by another version, so
holding two releases side by side is the normal state rather than a
special case.
The bytes underneath are shared. A file lives once, at
blobs/sha256/, under its own checksum, and the version slot
holds a name for it: a hardlink where the filesystem allows one, a
symlink or a copy where it does not. Two packages declaring the same 4
GB file therefore keep one copy and two independent dependency records.
They also transfer it once, because the lock is keyed on the checksum,
so the second session waits for the first and then finds the bytes
already there.
A resource composed from part() records stores its
pieces in the same place. Each part is admitted under its own digest, so
a base every version of a series declares is one blob however many
versions declare it, and publishing a version costs a consumer the delta
rather than the whole artefact. No version slot names a part. What
reaches it is the entry composed from it, which records the part digests
beside its provenance, and that is what keeps a base alive for exactly
as long as the last cached version still holding it.
Everything the cache owns is read-only. A caller writing to a
returned path would otherwise damage every package that shares those
bytes, so the write fails at the point of the mistake instead. A caller
that needs a writable layout declares a processor(), which
gets its own slot.
The store keeps no metadata of its own. Whether a blob is still needed is answered by reading the package indexes, so there is no reference count that a crash, a restored backup or a hand-deleted directory could leave disagreeing with them.
The processed result of a processor sits beside the raw artefact rather than replacing it, under a directory named for the processor id. Changing the transformation means changing the id, which invalidates the derived tree without touching the download it came from.
index.rds is one small file per package holding the
provenance records. It is written to a sibling temporary file and
renamed, so a reader never observes a half-written index. A metadata
database was considered and left out: the volume is tiny, and an atomic
per-package file removes a dependency and a class of locking
problems.
What a cached copy has to prove
Three different questions, three answers, and the entry record keeps them apart so that “verified” never quietly means “we looked at this sometime”.
| what it does | when it runs | recorded as | |
|---|---|---|---|
| full verification | re-hashes the bytes | on download, on verify = TRUE, and once the last one is
older than getaca.verify_days
|
verified_at |
| cheap check | compares size against the entry | on every ordinary access | checked_at |
| access | none | on every ordinary access | accessed_at |
The cheap check catches truncation, replacement by a different-sized
file, and most accidental edits, for the cost of a
file.info() call on a four-gigabyte file. A same-size
substitution passes it, which is the case the scheduled re-hash exists
to catch.
getaca("backbone", package = "yourpkg") # cheap check
getaca("backbone", package = "yourpkg", verify = TRUE) # full re-hash first
options(getaca.verify_days = 30) # re-hash more oftenA cached copy that fails either check raises
getaca_error_cache_corrupt rather than being silently
refetched, and the message names the clean-up call. Silent repair would
hide a disk going bad, and hide a colleague who edited a file in the
cache directory.
The failure is a verdict on bytes rather than on the slot that found
it. Bytes live once and every package declaring them holds its own
record, so a mismatch in the shared copy withdraws
verified_at from every other slot naming it and each
re-hashes on next access. A slot holding its own copy, which is what a
filesystem refusing links leaves, and a processed tree derived from the
bytes are answerable only for themselves.
A part is re-hashed every time a composition reuses it, since nothing else ever looks at it. Scheduled re-verification is driven from the entries, and a part blob is named by an entry’s part list rather than being an entry of its own, so a base that rotted on disk would otherwise surface as the declaration failing to produce its own artefact. A part failing its own checksum is dropped and fetched again the way a stale partial transfer is: the declaration lists mirrors for those bytes, and nothing else names them.
The fourth timestamp is fetched_at, which never moves.
Together the four answer questions that collapsing them would destroy: a
resource fetched in January, re-hashed in April and read this morning
reports exactly that.
Provenance
getaca_info("backbone", package = "yourpkg")
#> <getaca cache entry> yourpkg/backbone@2026-06
#> path ~/.cache/R/getaca/yourpkg/backbone/2026-06/raw/backbone-2026-06.zip
#> sha256 9f9f9f...
#> size 797,000,000 bytes
#> license CC-BY-4.0
#> built from source_release: 2026-06
#> resolved by current registry sha256:8b31e0da54cf (published 2026-07-22)
#> source url https://host.invalid/backbone-2026-06.zip
#> getaca 0.0.0.9000
#> fetched 2026-07-26 11:02:13
#> verified 2026-07-26 11:09:44 (full re-hash)
#> checked 2026-07-26 15:31:02 (size and mtime)A composed resource reports what produced it, under the source url its series leaves empty:
#> source url NA
#> composed 3 parts via 'concat'
#> part 919191919191
#> part 4e4e4e4e4e4e
#> part 777777777777An uncached resource gives NULL, which is what makes the
call safe in a report covering a machine that holds some of the set:
is.null(getaca_info("backbone", registry = reg))
#> [1] TRUEgetaca_catalogue() is the same information across
everything, plus the declarations that have never been downloaded:
getaca_catalogue(registry = reg)[, c("package", "name", "version",
"current", "declared", "cached")]
#> package name version current declared cached
#> 1 yourpkg backbone 2026-06 TRUE TRUE FALSEWith no arguments it covers every installed package that ships a registry together with every package holding cached resources, which is the report worth pasting into an issue:
str(getaca_catalogue(), max.level = 1)
#> 'data.frame': 0 obs. of 18 variables:
#> $ package : chr
#> $ name : chr
#> $ version : chr
#> $ current : logi
#> $ processor : chr
#> $ parts : int
#> $ link : chr
#> $ declared : logi
#> $ cached : logi
#> $ size : num
#> $ license : chr
#> $ doi : chr
#> $ source : chr
#> $ registry_digest: chr
#> $ verified_at : 'POSIXct' num
#> $ accessed_at : 'POSIXct' num
#> $ pinned : logi
#> $ path : chrThe columns worth knowing: size in bytes,
license, source and
registry_digest naming the policy and the registry state
that resolved it, parts for how many pieces the artefact
was composed from and 0 where it was served whole, the
three timestamps, pinned, and path.
Two sessions, one download
Two R sessions asking for the same four-gigabyte file must not both fetch it, and must never mistake each other’s in-flight temporary file for a finished resource.
The lock is a directory under .locks/, named for the
declared checksum. dir.create() is atomic on both POSIX and
Windows, which makes a directory a portable mutex with no compiled
dependency and no lockfile library.
Keying it on the checksum rather than on the resource triple means the two sessions need not be asking on behalf of the same package. Two packages declaring the same file are waiting for the same transfer, and the one that waits finds the bytes in the store when it wakes.
What a second session does:
tries to create the lock directory, and fails
checks whether the lock is stale, by the age of the holder file inside it
waits, polling, until the holder releases
re-reads the cache index, finds the first session’s entry, and returns that path
Step 4 is the point. The waiter reads the entry the first session wrote and returns that path, so the second transfer never starts. The cache check is repeated after the lock is acquired precisely because the situation may have changed while waiting.
Composing from parts takes one lock per distinct part digest, held
for the whole composition rather than for each transfer. Two sessions
must not both fetch one part, which is what the acquisition lock already
does for a whole file. And a part blob is named by no index until the
entry composed from it is written, so the lock is also what tells
another session that bytes nothing references yet are wanted: the
unreferenced sweep treats a blob under an active lock as
live.
A lock whose holder died leaves a directory behind. It goes stale
after getaca.lock_stale_seconds, defaulting to 1800, after
which the next session removes it and takes over. A session that waits
longer than its timeout gets an error naming the lock path and the
unlink() call that clears it, so a genuinely wedged lock is
a one-line fix rather than a support thread.
options(getaca.lock_stale_seconds = 600)Set it lower for short downloads on a shared machine, higher when a single transfer legitimately runs for an hour.
How bytes get in
Transfers land in .tmp/, are sized, hashed, and only
then moved into place. An interrupted transfer can never appear as a
valid cached resource, and a failed transfer never touches a copy that
was already good.
The temporary file is named after the declared checksum and the mirror that produced it. Naming it after the checksum makes an interrupted download resumable on the next attempt, which matters when the resource is measured in gigabytes. Giving each mirror its own file matters for a subtler reason: a partial transfer is resumable only against the host that produced it, so sharing one file across mirrors would let a failed attempt at the first be resumed onto by the second, and the resulting corruption is indistinguishable from the publisher having changed the bytes.
A resumed transfer that completes but does not verify indicts the partial file rather than the publisher, so the same mirror is asked once more from empty before any conclusion is drawn about upstream. Without that, one stale temporary file makes a resource permanently unfetchable and blames the wrong party for it.
Verified bytes are then admitted to the store under their own
checksum, and the version slot is given a name for them. Admission is a
rename, since .tmp/ and blobs/ share the cache
root and therefore share a filesystem. Bytes already in the store are
already named by their checksum, so admitting the same file a second
time is a no-op and the temporary copy is dropped.
A composed resource joins at that same point. Each part lands in
.tmp/ and is admitted under its own digest, the series is
combined into a second temporary file named after the artefact’s
checksum, and that file is hashed against the record before anything
else sees it. Only a result matching the declaration is admitted and
given its view, so a combiner() needs no more trust than a
mirror does: it cannot produce bytes the declaration did not already
name. A part already in the store is used where it lies, which is the
whole point of declaring one, and an interrupted composition is
overwritten by the retry rather than accumulating.
Watching a transfer
getaca drives its own transfer loop, so what a download looks like is a setting rather than whatever the transfer library prints.
getaca_progress("bar") # redraws one line, the default when interactive
getaca_progress("line") # one line to start and one to finish, for a log
getaca_progress("none") # nothingyourpkg/backbone@2026-09 [============> ] 63% 512 MB / 812 MB 41 MB/s ETA 00:07
The share is measured against the size the registry declares, which is known before the first byte arrives and stays right when a mirror sends no content length. A resource composed from parts reports each piece under its own label, so a series reads as one download in stages:
yourpkg/backbone@2026-09 (part 1 of 3) [===================] 100% 797 MB in 00:19
yourpkg/backbone@2026-09 (part 2 of 3) [========> ] 44% 4.0 MB / 9.1 MB ...
quiet = TRUE on a single call reports nothing whatever
the session is set to, so one silent retrieval never needs the setting
changed and put back:
getaca("backbone", package = "yourpkg", quiet = TRUE)A package that wants a download to look like its own writes a
reporter(), which is a function of one argument:
counter <- reporter("counter", function(event) {
if (identical(event$type, "end") && identical(event$status, "ok")) {
message(format(event$id), ": ", event$bytes, " bytes")
}
})
counter
#> <getaca reporter> counterThe events are begin, bytes and
end, and a handler switches on event$type and
ignores what it does not use. begin carries the resource,
the mirror, the declared total and the offset
an interrupted transfer resumed from; bytes carries the
cumulative count; end carries the outcome. See
?"getaca-progress" for the fields.
A reporter never decides whether a retrieval succeeds. One that raises is caught, reported once as a warning, and switched off for the rest of the call.
Retention
CRAN permits tools::R_user_dir() on condition that
contents are “actively managed (including removing outdated material)”.
getaca reads that as a retention policy rather than a
function users might discover, so collection runs automatically after
every successful retrieval.
Five sweeps, cheapest and safest first:
| Sweep | Removes | Clock |
|---|---|---|
broken |
entries whose path is missing or fails the cheap check | none |
temp |
abandoned transfers in .tmp/
|
7 days by mtime |
superseded |
unpinned versions the registry no longer names |
getaca.supersede_days, default 30 |
lru |
least recently used unpinned entries | only above getaca.max_bytes, default 20 GB |
unreferenced |
bytes in the store that no entry names any more | none |
The first four sweeps remove names. unreferenced runs
last and removes the bytes those names were for, once the last one is
gone. A blob under an active lock is left alone: it belongs to a session
that has admitted it and has not yet written its entry.
What an entry names, for this purpose, is the blob it holds together with the parts that blob was composed from. Reachability runs over both, so a base stays for as long as some cached version is still composed from it and goes with the last one.
The size ceiling measures what the cache occupies, so shared bytes
count once. Two packages declaring the same 4 GB file count 4 GB against
getaca.max_bytes, and evicting one of them frees nothing
until the other goes too. Part blobs are occupancy on the same terms:
evicting one version of a series reclaims the pieces no surviving
version holds, and leaves a base that another version still
declares.
Superseded and not-recently-used age on separate clocks on purpose. An expensive resource that is still the current version is never dropped merely for being old; it is dropped only when the cache is over its ceiling, and then only after everything already useless has gone.
Three things are never removed: pinned entries, the version the registry currently names, and anything under an active lock.
The automatic pass after a retrieval runs only broken
and temp, the two sweeps that can only ever remove material
which is already useless. Reclaiming a superseded four-gigabyte version
is a decision, so it happens on a schedule rather than as a side effect
of a download.
Cleaning by hand
getaca_clean(dry_run = TRUE)
#> [1] package resource reason bytes path
#> <0 rows> (or 0-length row.names)An empty result on a fresh cache. On a working one, each row names the package, the resource, the reason and the bytes it would reclaim, which is the report to read before running it for real.
getaca_clean(dry_run = TRUE)
#> package resource reason bytes
#> 1 yourpkg yourpkg/backbone@2026-03 superseded version past... 7.97e+08
#> 2 <NA> <NA> abandoned transfer 1.20e+07
getaca_clean() # run every sweep
getaca_clean(what = "temp") # just the abandoned transfers
getaca_clean(package = "yourpkg") # one package
getaca_clean("backbone", package = "yourpkg") # one resource nameKeeping something the sweeps would otherwise take:
getaca_keep("backbone", package = "yourpkg") # pin it
getaca_keep("backbone", package = "yourpkg", pinned = FALSE) # release the pinA pinned entry is exempt from the superseded and LRU sweeps
permanently. It is still subject to the broken sweep,
because an entry whose bytes are gone is not worth protecting.
Settings
Every setting is readable from an option or an environment variable, with the option taking precedence.
| Option | Environment variable | Default | Controls |
|---|---|---|---|
getaca.cache |
GETACA_CACHE |
R_user_dir() |
where everything lives |
getaca.policy |
GETACA_POLICY |
registry default | which channel resolves |
getaca.progress |
GETACA_PROGRESS |
auto |
what a transfer looks like |
getaca.verify_days |
GETACA_VERIFY_DAYS |
90 | scheduled re-hash interval |
getaca.supersede_days |
GETACA_SUPERSEDE_DAYS |
30 | retention for undeclared versions |
getaca.max_bytes |
GETACA_MAX_BYTES |
20 GB | ceiling above which LRU runs |
getaca.timeout |
GETACA_TIMEOUT |
3600 | transfer timeout in seconds |
getaca.lock_stale_seconds |
GETACA_LOCK_STALE_SECONDS |
1800 | when a lock is abandoned |
getaca.pin_file |
getaca.pins.rds in the working directory |
where pins are read from |
The defaults suit a laptop holding a couple of large reference
datasets. Two are worth revisiting on a shared machine: raise
getaca.max_bytes when the cache lives on a volume sized for
it, and lower getaca.lock_stale_seconds when transfers are
short and a wedged lock costs more than a rare duplicate download.
options(
getaca.cache = "/mnt/data/getaca",
getaca.max_bytes = 200 * 1024^3,
getaca.verify_days = 30
)Moving a cache
Copy the directory. There are no absolute paths recorded inside it, so a cache built on one machine works on another:
# on a connected machine
Sys.setenv(GETACA_CACHE = "/tmp/seed")
getaca_prefetch(package = "yourpkg")
# then, on the machine that has no network
Sys.setenv(GETACA_CACHE = "/opt/getaca")
getaca_catalogue() # everything already thereThe same property is what makes CI caching work: the archive an actions cache restores is the cache, with nothing to rebuild.
Where to go next
vignette("checks")for seeding a cache in CIvignette("failures")forgetaca_error_cache_corruptand its neighboursvignette("policies")for what decides which version lands in the cachevignette("declaring")for declaring a resource that arrives as a series