Skip to contents

The cache is where a declared resource becomes a local path. This article covers the layout, what makes a cached copy trustworthy, how two sessions avoid downloading the same file twice, and the retention policy CRAN requires of a package that writes to a user directory.

None of it is private. The cache is an ordinary directory tree, and that is a deliberate property: copying it to an offline machine, restoring it from a CI cache action, or looking at it with a file browser all work.

Where it lives

getaca_cache_dir()
#> [1] "<tempdir>/getaca-cache-vignette"

That is the sandbox this vignette runs in. The default is tools::R_user_dir("getaca", "cache"), the location CRAN permits for cached downloads. Two overrides take precedence, in this order:

options(getaca.cache = "/mnt/fast/getaca")   # this session
Sys.setenv(GETACA_CACHE = "/mnt/fast/getaca") # this process and its children

The environment variable is the one to reach for in CI and in job scripts, because it survives into the R processes a build step spawns. Asking for the directory does not create it; the first successful retrieval does.

Layout

<cache>/
  blobs/sha256/<aa>/<sha256>     verified bytes, named by their own checksum
  .locks/                        one per checksum, held during a transfer
  .tmp/                          in-flight downloads, never visible as cache
  <package>/
    index.rds                    provenance for this package only
    <name>/<version>/
      raw/<file>                 this slot's name for a blob
      proc-<processor-id>/       processed result, own provenance

Everything a package declares is scoped by declaring package, then resource name, then version. That falls out of identity being the triple package / name / version, and it buys two things. Two packages declaring a resource called "backbone" never share a slot, so one package’s registry update cannot affect another’s cached data. And a version can never be overwritten by another version, so holding two releases side by side is the normal state rather than a special case.

The bytes underneath are shared. A file lives once, at blobs/sha256/, under its own checksum, and the version slot holds a name for it: a hardlink where the filesystem allows one, a symlink or a copy where it does not. Two packages declaring the same 4 GB file therefore keep one copy and two independent dependency records. They also transfer it once, because the lock is keyed on the checksum, so the second session waits for the first and then finds the bytes already there.

A resource composed from part() records stores its pieces in the same place. Each part is admitted under its own digest, so a base every version of a series declares is one blob however many versions declare it, and publishing a version costs a consumer the delta rather than the whole artefact. No version slot names a part. What reaches it is the entry composed from it, which records the part digests beside its provenance, and that is what keeps a base alive for exactly as long as the last cached version still holding it.

Everything the cache owns is read-only. A caller writing to a returned path would otherwise damage every package that shares those bytes, so the write fails at the point of the mistake instead. A caller that needs a writable layout declares a processor(), which gets its own slot.

The store keeps no metadata of its own. Whether a blob is still needed is answered by reading the package indexes, so there is no reference count that a crash, a restored backup or a hand-deleted directory could leave disagreeing with them.

The processed result of a processor sits beside the raw artefact rather than replacing it, under a directory named for the processor id. Changing the transformation means changing the id, which invalidates the derived tree without touching the download it came from.

index.rds is one small file per package holding the provenance records. It is written to a sibling temporary file and renamed, so a reader never observes a half-written index. A metadata database was considered and left out: the volume is tiny, and an atomic per-package file removes a dependency and a class of locking problems.

What a cached copy has to prove

Three different questions, three answers, and the entry record keeps them apart so that “verified” never quietly means “we looked at this sometime”.

what it does when it runs recorded as
full verification re-hashes the bytes on download, on verify = TRUE, and once the last one is older than getaca.verify_days verified_at
cheap check compares size against the entry on every ordinary access checked_at
access none on every ordinary access accessed_at

The cheap check catches truncation, replacement by a different-sized file, and most accidental edits, for the cost of a file.info() call on a four-gigabyte file. A same-size substitution passes it, which is the case the scheduled re-hash exists to catch.

getaca("backbone", package = "yourpkg")                 # cheap check
getaca("backbone", package = "yourpkg", verify = TRUE)  # full re-hash first
options(getaca.verify_days = 30)                  # re-hash more often

A cached copy that fails either check raises getaca_error_cache_corrupt rather than being silently refetched, and the message names the clean-up call. Silent repair would hide a disk going bad, and hide a colleague who edited a file in the cache directory.

The failure is a verdict on bytes rather than on the slot that found it. Bytes live once and every package declaring them holds its own record, so a mismatch in the shared copy withdraws verified_at from every other slot naming it and each re-hashes on next access. A slot holding its own copy, which is what a filesystem refusing links leaves, and a processed tree derived from the bytes are answerable only for themselves.

A part is re-hashed every time a composition reuses it, since nothing else ever looks at it. Scheduled re-verification is driven from the entries, and a part blob is named by an entry’s part list rather than being an entry of its own, so a base that rotted on disk would otherwise surface as the declaration failing to produce its own artefact. A part failing its own checksum is dropped and fetched again the way a stale partial transfer is: the declaration lists mirrors for those bytes, and nothing else names them.

The fourth timestamp is fetched_at, which never moves. Together the four answer questions that collapsing them would destroy: a resource fetched in January, re-hashed in April and read this morning reports exactly that.

Provenance

getaca_info("backbone", package = "yourpkg")
#> <getaca cache entry> yourpkg/backbone@2026-06
#>   path        ~/.cache/R/getaca/yourpkg/backbone/2026-06/raw/backbone-2026-06.zip
#>   sha256      9f9f9f...
#>   size        797,000,000 bytes
#>   license     CC-BY-4.0
#>   built from  source_release: 2026-06
#>   resolved by current registry sha256:8b31e0da54cf (published 2026-07-22)
#>   source url  https://host.invalid/backbone-2026-06.zip
#>   getaca      0.0.0.9000
#>   fetched     2026-07-26 11:02:13
#>   verified    2026-07-26 11:09:44 (full re-hash)
#>   checked     2026-07-26 15:31:02 (size and mtime)

A composed resource reports what produced it, under the source url its series leaves empty:

#>   source url  NA
#>   composed    3 parts via 'concat'
#>     part      919191919191
#>     part      4e4e4e4e4e4e
#>     part      777777777777

An uncached resource gives NULL, which is what makes the call safe in a report covering a machine that holds some of the set:

is.null(getaca_info("backbone", registry = reg))
#> [1] TRUE

getaca_catalogue() is the same information across everything, plus the declarations that have never been downloaded:

getaca_catalogue(registry = reg)[, c("package", "name", "version",
                                     "current", "declared", "cached")]
#>   package     name version current declared cached
#> 1 yourpkg backbone 2026-06    TRUE     TRUE  FALSE

With no arguments it covers every installed package that ships a registry together with every package holding cached resources, which is the report worth pasting into an issue:

str(getaca_catalogue(), max.level = 1)
#> 'data.frame':    0 obs. of  18 variables:
#>  $ package        : chr 
#>  $ name           : chr 
#>  $ version        : chr 
#>  $ current        : logi 
#>  $ processor      : chr 
#>  $ parts          : int 
#>  $ link           : chr 
#>  $ declared       : logi 
#>  $ cached         : logi 
#>  $ size           : num 
#>  $ license        : chr 
#>  $ doi            : chr 
#>  $ source         : chr 
#>  $ registry_digest: chr 
#>  $ verified_at    : 'POSIXct' num 
#>  $ accessed_at    : 'POSIXct' num 
#>  $ pinned         : logi 
#>  $ path           : chr

The columns worth knowing: size in bytes, license, source and registry_digest naming the policy and the registry state that resolved it, parts for how many pieces the artefact was composed from and 0 where it was served whole, the three timestamps, pinned, and path.

Two sessions, one download

Two R sessions asking for the same four-gigabyte file must not both fetch it, and must never mistake each other’s in-flight temporary file for a finished resource.

The lock is a directory under .locks/, named for the declared checksum. dir.create() is atomic on both POSIX and Windows, which makes a directory a portable mutex with no compiled dependency and no lockfile library.

Keying it on the checksum rather than on the resource triple means the two sessions need not be asking on behalf of the same package. Two packages declaring the same file are waiting for the same transfer, and the one that waits finds the bytes in the store when it wakes.

What a second session does:

  1. tries to create the lock directory, and fails

  2. checks whether the lock is stale, by the age of the holder file inside it

  3. waits, polling, until the holder releases

  4. re-reads the cache index, finds the first session’s entry, and returns that path

Step 4 is the point. The waiter reads the entry the first session wrote and returns that path, so the second transfer never starts. The cache check is repeated after the lock is acquired precisely because the situation may have changed while waiting.

Composing from parts takes one lock per distinct part digest, held for the whole composition rather than for each transfer. Two sessions must not both fetch one part, which is what the acquisition lock already does for a whole file. And a part blob is named by no index until the entry composed from it is written, so the lock is also what tells another session that bytes nothing references yet are wanted: the unreferenced sweep treats a blob under an active lock as live.

A lock whose holder died leaves a directory behind. It goes stale after getaca.lock_stale_seconds, defaulting to 1800, after which the next session removes it and takes over. A session that waits longer than its timeout gets an error naming the lock path and the unlink() call that clears it, so a genuinely wedged lock is a one-line fix rather than a support thread.

options(getaca.lock_stale_seconds = 600)

Set it lower for short downloads on a shared machine, higher when a single transfer legitimately runs for an hour.

How bytes get in

Transfers land in .tmp/, are sized, hashed, and only then moved into place. An interrupted transfer can never appear as a valid cached resource, and a failed transfer never touches a copy that was already good.

The temporary file is named after the declared checksum and the mirror that produced it. Naming it after the checksum makes an interrupted download resumable on the next attempt, which matters when the resource is measured in gigabytes. Giving each mirror its own file matters for a subtler reason: a partial transfer is resumable only against the host that produced it, so sharing one file across mirrors would let a failed attempt at the first be resumed onto by the second, and the resulting corruption is indistinguishable from the publisher having changed the bytes.

A resumed transfer that completes but does not verify indicts the partial file rather than the publisher, so the same mirror is asked once more from empty before any conclusion is drawn about upstream. Without that, one stale temporary file makes a resource permanently unfetchable and blames the wrong party for it.

Verified bytes are then admitted to the store under their own checksum, and the version slot is given a name for them. Admission is a rename, since .tmp/ and blobs/ share the cache root and therefore share a filesystem. Bytes already in the store are already named by their checksum, so admitting the same file a second time is a no-op and the temporary copy is dropped.

A composed resource joins at that same point. Each part lands in .tmp/ and is admitted under its own digest, the series is combined into a second temporary file named after the artefact’s checksum, and that file is hashed against the record before anything else sees it. Only a result matching the declaration is admitted and given its view, so a combiner() needs no more trust than a mirror does: it cannot produce bytes the declaration did not already name. A part already in the store is used where it lies, which is the whole point of declaring one, and an interrupted composition is overwritten by the retry rather than accumulating.

Watching a transfer

getaca drives its own transfer loop, so what a download looks like is a setting rather than whatever the transfer library prints.

getaca_progress("bar")    # redraws one line, the default when interactive
getaca_progress("line")   # one line to start and one to finish, for a log
getaca_progress("none")   # nothing
yourpkg/backbone@2026-09  [============>      ]  63%  512 MB / 812 MB  41 MB/s  ETA 00:07

The share is measured against the size the registry declares, which is known before the first byte arrives and stays right when a mirror sends no content length. A resource composed from parts reports each piece under its own label, so a series reads as one download in stages:

yourpkg/backbone@2026-09 (part 1 of 3)  [===================] 100%  797 MB in 00:19
yourpkg/backbone@2026-09 (part 2 of 3)  [========>          ]  44%  4.0 MB / 9.1 MB  ...

quiet = TRUE on a single call reports nothing whatever the session is set to, so one silent retrieval never needs the setting changed and put back:

getaca("backbone", package = "yourpkg", quiet = TRUE)

A package that wants a download to look like its own writes a reporter(), which is a function of one argument:

counter <- reporter("counter", function(event) {
  if (identical(event$type, "end") && identical(event$status, "ok")) {
    message(format(event$id), ": ", event$bytes, " bytes")
  }
})
counter
#> <getaca reporter> counter

The events are begin, bytes and end, and a handler switches on event$type and ignores what it does not use. begin carries the resource, the mirror, the declared total and the offset an interrupted transfer resumed from; bytes carries the cumulative count; end carries the outcome. See ?"getaca-progress" for the fields.

A reporter never decides whether a retrieval succeeds. One that raises is caught, reported once as a warning, and switched off for the rest of the call.

Retention

CRAN permits tools::R_user_dir() on condition that contents are “actively managed (including removing outdated material)”. getaca reads that as a retention policy rather than a function users might discover, so collection runs automatically after every successful retrieval.

Five sweeps, cheapest and safest first:

Sweep Removes Clock
broken entries whose path is missing or fails the cheap check none
temp abandoned transfers in .tmp/ 7 days by mtime
superseded unpinned versions the registry no longer names getaca.supersede_days, default 30
lru least recently used unpinned entries only above getaca.max_bytes, default 20 GB
unreferenced bytes in the store that no entry names any more none

The first four sweeps remove names. unreferenced runs last and removes the bytes those names were for, once the last one is gone. A blob under an active lock is left alone: it belongs to a session that has admitted it and has not yet written its entry.

What an entry names, for this purpose, is the blob it holds together with the parts that blob was composed from. Reachability runs over both, so a base stays for as long as some cached version is still composed from it and goes with the last one.

The size ceiling measures what the cache occupies, so shared bytes count once. Two packages declaring the same 4 GB file count 4 GB against getaca.max_bytes, and evicting one of them frees nothing until the other goes too. Part blobs are occupancy on the same terms: evicting one version of a series reclaims the pieces no surviving version holds, and leaves a base that another version still declares.

Superseded and not-recently-used age on separate clocks on purpose. An expensive resource that is still the current version is never dropped merely for being old; it is dropped only when the cache is over its ceiling, and then only after everything already useless has gone.

Three things are never removed: pinned entries, the version the registry currently names, and anything under an active lock.

The automatic pass after a retrieval runs only broken and temp, the two sweeps that can only ever remove material which is already useless. Reclaiming a superseded four-gigabyte version is a decision, so it happens on a schedule rather than as a side effect of a download.

Cleaning by hand

getaca_clean(dry_run = TRUE)
#> [1] package  resource reason   bytes    path    
#> <0 rows> (or 0-length row.names)

An empty result on a fresh cache. On a working one, each row names the package, the resource, the reason and the bytes it would reclaim, which is the report to read before running it for real.

getaca_clean(dry_run = TRUE)
#>   package                 resource                     reason    bytes
#> 1 yourpkg yourpkg/backbone@2026-03 superseded version past... 7.97e+08
#> 2    <NA>                     <NA>         abandoned transfer 1.20e+07

getaca_clean()                                 # run every sweep
getaca_clean(what = "temp")                    # just the abandoned transfers
getaca_clean(package = "yourpkg")              # one package
getaca_clean("backbone", package = "yourpkg")  # one resource name

Keeping something the sweeps would otherwise take:

getaca_keep("backbone", package = "yourpkg")                  # pin it
getaca_keep("backbone", package = "yourpkg", pinned = FALSE)  # release the pin

A pinned entry is exempt from the superseded and LRU sweeps permanently. It is still subject to the broken sweep, because an entry whose bytes are gone is not worth protecting.

Settings

Every setting is readable from an option or an environment variable, with the option taking precedence.

Option Environment variable Default Controls
getaca.cache GETACA_CACHE R_user_dir() where everything lives
getaca.policy GETACA_POLICY registry default which channel resolves
getaca.progress GETACA_PROGRESS auto what a transfer looks like
getaca.verify_days GETACA_VERIFY_DAYS 90 scheduled re-hash interval
getaca.supersede_days GETACA_SUPERSEDE_DAYS 30 retention for undeclared versions
getaca.max_bytes GETACA_MAX_BYTES 20 GB ceiling above which LRU runs
getaca.timeout GETACA_TIMEOUT 3600 transfer timeout in seconds
getaca.lock_stale_seconds GETACA_LOCK_STALE_SECONDS 1800 when a lock is abandoned
getaca.pin_file getaca.pins.rds in the working directory where pins are read from

The defaults suit a laptop holding a couple of large reference datasets. Two are worth revisiting on a shared machine: raise getaca.max_bytes when the cache lives on a volume sized for it, and lower getaca.lock_stale_seconds when transfers are short and a wedged lock costs more than a rare duplicate download.

options(
  getaca.cache        = "/mnt/data/getaca",
  getaca.max_bytes    = 200 * 1024^3,
  getaca.verify_days  = 30
)

Moving a cache

Copy the directory. There are no absolute paths recorded inside it, so a cache built on one machine works on another:

# on a connected machine
Sys.setenv(GETACA_CACHE = "/tmp/seed")
getaca_prefetch(package = "yourpkg")

# then, on the machine that has no network
Sys.setenv(GETACA_CACHE = "/opt/getaca")
getaca_catalogue()   # everything already there

The same property is what makes CI caching work: the archive an actions cache restores is the cache, with nothing to rebuild.

Where to go next