Skip to contents

Readings in long form to a [unit, bin, channel] array, reachable without the fitting layer.

All of the Python reference

grain_matrix()

grain_matrix(
    data=None,
    id=None,
    time=None,
    value=None,
    *,
    grain='day',
    stats=('mean',),
    year_start='01-01',
    partial='keep',
    tz=None,
)

Bin readings by the calendar and summarise every bin.

data is a mapping of column name to sequence, or any object with __getitem__ over the three column names given by id, time and value. Naming two or more grains returns a TimesiftSet; naming one, whether as a string or as a sequence of one, returns the representation itself. grain may also be a callable, which is handed the reading instants and must return the start of each reading’s bin.

tz names the calendar to bin by. Left at None the instants are taken as already expressed in that calendar, which is what a zone-free datetime64 says and what the R side does for a series carried in UTC. Given a zone name, the instants are read as UTC and binned by that zone’s clock, which is what the R side does for a series carrying a tzone: the same instants and the same zone give the same answer in both languages. A time column that carries a zone of its own names the calendar the same way, so a zone-aware column bins by its own clock without being told to; naming a different one in tz beside it is an error. The "native" grain is the one grain not read on that clock: its bin is the reading itself, so the two readings of an hour a zone repeats are two bins, and the record read at "native" is the same array whichever zone it is carried in.

partial says what becomes of a bin the record does not cover for its whole calendar span, which is what a record beginning or ending away from a bin boundary produces. "keep", the default, returns it alongside the full bins; "drop" removes it. Either way the verdict is carried on bin_partial, so a kept partial bin is labelled rather than silent. A caller-supplied binning declares its own bins, so the package cannot know where the last one was meant to end and takes the record’s end as its end.

lookback_matrix()

lookback_matrix(
    data=None,
    id=None,
    time=None,
    value=None,
    at=None,
    span=None,
    *,
    lag='0 days',
    bins=1,
    stats='mean',
    tz=None,
)

Read a fixed length of record ending a fixed lag before each target’s own instant.

It is the reduction a calendar cannot express: two targets on the same unit a fortnight apart read two different stretches of the same series, so the bins are relative to the target rather than to a month or a week.

at is a mapping with an "id" array of units and an "at" array of anchor instants, one row per target; a unit may carry any number of them. Bin b of a target anchored at a covers [a - lag - span + b * step, a - lag - span + (b + 1) * step), with step the span divided by bins and b counted from zero. Only the readings of the target’s own unit are read, and every (target, bin) cell must hold at least one: a lookback reaching past the record is an error naming the target, never a padded row.

span and lag are read from a count and a unit – "30 days", "12 hours", "1 year" – or from a bare number of seconds. A year is 365 days and a month is 30 days here, because a lookback of a fixed length is a fixed length rather than a calendar step.

tz names the calendar, as it does for grain_matrix. The anchors are instants and are read as a clock in that same calendar, so one record is binned by one calendar. The span is measured on that clock: a lookback of one day ending at a local midnight holds the whole local day before it, which is 25 hours of record on the night a zone sets its clock back and 23 on the night it sets it forward. That is what keeps a calendar day whole inside a bin for the four day-level statistics; a length fixed in instants could not.

coverage()

coverage(data=None, id=None, time=None, *, grain='day', year_start='01-01', tz=None)

Which units reach which bins.

A representation needs every unit in every bin, and grain_matrix refuses a record where one is missing rather than pad it. This is the same binning laid out so the gaps can be read: how many readings each unit has in each bin, over every bin the calendar tiles the record with from the first bin any unit touches to the last. What to do about a gap is the analyst’s decision, and this is the table it is made on; nothing here fills a cell. grain and tz read as they do for grain_matrix.

Coverage

Coverage(count, units, bins, grain, bin_start)

How many readings each unit has in each bin, over every bin the calendar tiles the record with. count is [unit, bin]; a unit that started late, stopped early or lost a month is a row with zeros in it, and a bin the whole record skips is a column of zeros.

Attributes:

  • count - np.ndarray
  • units - tuple[str, …]
  • bins - tuple[str, …]
  • grain - str
  • bin_start - np.ndarray

empty

The [unit, bin] mask of cells holding no reading.

units_with_gaps()

units_with_gaps(self)

The units that do not reach every bin.

bins_no_unit_reaches()

bins_no_unit_reaches(self)

The bins the whole record skips.

timesift_set()

timesift_set(x)

Every entry point that fits across grains takes a representation, a set, or a bare mapping, and works on a set. One coercion, so no caller repeats the three cases.

calendar_channels()

calendar_channels(x: TimesiftMatrix, cycles=('year',))

Where in the year, or the day, each bin sits, as the sine and cosine of its fractional position in each cycle named, in the order given.

The position is read at the midpoint of the record each bin holds, in UTC. The day cycle reads bins that sit less than a day apart and is refused otherwise. inst/spec/representation.md is the normative description.

bind_channels()

bind_channels(*parts: TimesiftMatrix)

Put the channels of several representations of the same units and bins side by side.

The channels come back in the order the arguments are given and, inside each argument, in its own channel order. Everything else is the first argument’s, except static and position, which name channels and so name those of every argument.

feature_matrix()

feature_matrix(m, units=None, features=None, label: str = 'features')

Bring an already-reduced feature table into a ladder as a one-channel representation.

It carries no time axis, because it has none: the reduction already happened, elsewhere, and what reaches the model is a list of numbers per unit. That is the whole point of comparing against it.

TimesiftMatrix

TimesiftMatrix(
    values,
    units,
    bins,
    stats,
    grain,
    year_start,
    bin_start,
    bin_end,
    bin_n,
    bin_partial,
    span,
    lag,
    static,
    position,
    coords,
    unit_ids,
)

A [row, bin, channel] representation and the reduction that produced it.

A row is a unit where the calendar did the binning and a target where a lookback did, since a unit carrying several targets cannot name a row on its own. span and lag are set by a lookback alone, and are what rebuilding one for new targets reads. static names the channels holding the same number in every bin, which flatten reads once each. position names the channels calendar_channels made, which the encoders read at their own amplitude rather than standardise.

bin_start, bin_end and bin_partial are the calendar’s, and are None on a representation the calendar did not bin.

coords is each row’s pair of coordinates and unit_ids the identifier of the unit each row belongs to, where timesift.timesift was given coords and id. They place a row and are not channels, so no learner reads them as a predictor, and a split of the rows splits them with it.

Attributes:

  • values - np.ndarray
  • units - tuple[str, …]
  • bins - tuple[str, …]
  • stats - tuple[str, …]
  • grain - str
  • year_start - str | None
  • bin_start - np.ndarray | None
  • bin_end - np.ndarray | None
  • bin_n - np.ndarray
  • bin_partial - np.ndarray | None
  • span - int | None
  • lag - int | None
  • static - tuple[str, …]
  • position - tuple[str, …]
  • coords - np.ndarray | None
  • unit_ids - tuple[str, …] | None

shape

Rows, bins and channels.

channel()

channel(self, name: str)

One statistic as a [row, bin] matrix.

take_units()

take_units(self, index)

The representation restricted to a subset of its rows, in the order given.

TimesiftSet

A ladder of representations, one per grain.

Naming several grains in grain_matrix returns one of these: representations of the same units, differing only in how coarsely the record was read. It is what grain_ladder fits across, and it reads as a mapping of grain name to representation.

units

The units the set covers, which every grain in it shares.

GRAINS

GRAINS = ('native', 'halfday', 'day', 'week', 'month', 'season', 'year')

STATS

STATS = ('mean', 'min', 'max') + DAY_LEVEL_STATS

DAY_LEVEL_STATS

DAY_LEVEL_STATS = ('cold_day', 'warm_day', 'mean_daily_min', 'mean_daily_max')