Readings in long form to a [unit, bin, channel] array,
reachable without the fitting layer.
grain_matrix()
grain_matrix(
data=None,
id=None,
time=None,
value=None,
*,
grain='day',
stats=('mean',),
year_start='01-01',
partial='keep',
tz=None,
)Bin readings by the calendar and summarise every bin.
data is a mapping of column name to sequence, or any
object with __getitem__ over the three column names given
by id, time and value. Naming two
or more grains returns a TimesiftSet; naming one, whether
as a string or as a sequence of one, returns the representation itself.
grain may also be a callable, which is handed the reading
instants and must return the start of each reading’s bin.
tz names the calendar to bin by. Left at
None the instants are taken as already expressed in that
calendar, which is what a zone-free datetime64 says and
what the R side does for a series carried in UTC. Given a zone name, the
instants are read as UTC and binned by that zone’s clock, which is what
the R side does for a series carrying a tzone: the same
instants and the same zone give the same answer in both languages. A
time column that carries a zone of its own names the calendar the same
way, so a zone-aware column bins by its own clock without being told to;
naming a different one in tz beside it is an error. The
"native" grain is the one grain not read on that clock: its
bin is the reading itself, so the two readings of an hour a zone repeats
are two bins, and the record read at "native" is the same
array whichever zone it is carried in.
partial says what becomes of a bin the record does not
cover for its whole calendar span, which is what a record beginning or
ending away from a bin boundary produces. "keep", the
default, returns it alongside the full bins; "drop" removes
it. Either way the verdict is carried on bin_partial, so a
kept partial bin is labelled rather than silent. A caller-supplied
binning declares its own bins, so the package cannot know where the last
one was meant to end and takes the record’s end as its end.
lookback_matrix()
lookback_matrix(
data=None,
id=None,
time=None,
value=None,
at=None,
span=None,
*,
lag='0 days',
bins=1,
stats='mean',
tz=None,
)Read a fixed length of record ending a fixed lag before each target’s own instant.
It is the reduction a calendar cannot express: two targets on the same unit a fortnight apart read two different stretches of the same series, so the bins are relative to the target rather than to a month or a week.
at is a mapping with an "id" array of units
and an "at" array of anchor instants, one row per target; a
unit may carry any number of them. Bin b of a target
anchored at a covers
[a - lag - span + b * step, a - lag - span + (b + 1) * step),
with step the span divided by bins and
b counted from zero. Only the readings of the target’s own
unit are read, and every (target, bin) cell must hold at
least one: a lookback reaching past the record is an error naming the
target, never a padded row.
span and lag are read from a count and a
unit – "30 days", "12 hours",
"1 year" – or from a bare number of seconds. A year is 365
days and a month is 30 days here, because a lookback of a fixed length
is a fixed length rather than a calendar step.
tz names the calendar, as it does for
grain_matrix. The anchors are instants and are read as a
clock in that same calendar, so one record is binned by one calendar.
The span is measured on that clock: a lookback of one day ending at a
local midnight holds the whole local day before it, which is 25 hours of
record on the night a zone sets its clock back and 23 on the night it
sets it forward. That is what keeps a calendar day whole inside a bin
for the four day-level statistics; a length fixed in instants could
not.
coverage()
Which units reach which bins.
A representation needs every unit in every bin, and
grain_matrix refuses a record where one is missing rather
than pad it. This is the same binning laid out so the gaps can be read:
how many readings each unit has in each bin, over every bin the calendar
tiles the record with from the first bin any unit touches to the last.
What to do about a gap is the analyst’s decision, and this is the table
it is made on; nothing here fills a cell. grain and
tz read as they do for grain_matrix.
Coverage
How many readings each unit has in each bin, over every bin the
calendar tiles the record with. count is
[unit, bin]; a unit that started late, stopped early or
lost a month is a row with zeros in it, and a bin the whole record skips
is a column of zeros.
Attributes:
-
count- np.ndarray -
units- tuple[str, …] -
bins- tuple[str, …] -
grain- str -
bin_start- np.ndarray
timesift_set()
Every entry point that fits across grains takes a representation, a set, or a bare mapping, and works on a set. One coercion, so no caller repeats the three cases.
calendar_channels()
Where in the year, or the day, each bin sits, as the sine and cosine of its fractional position in each cycle named, in the order given.
The position is read at the midpoint of the record each bin holds, in
UTC. The day cycle reads bins that sit less than a day apart and is
refused otherwise. inst/spec/representation.md is the
normative description.
bind_channels()
Put the channels of several representations of the same units and bins side by side.
The channels come back in the order the arguments are given and,
inside each argument, in its own channel order. Everything else is the
first argument’s, except static and position,
which name channels and so name those of every argument.
feature_matrix()
Bring an already-reduced feature table into a ladder as a one-channel representation.
It carries no time axis, because it has none: the reduction already happened, elsewhere, and what reaches the model is a list of numbers per unit. That is the whole point of comparing against it.
TimesiftMatrix
TimesiftMatrix(
values,
units,
bins,
stats,
grain,
year_start,
bin_start,
bin_end,
bin_n,
bin_partial,
span,
lag,
static,
position,
coords,
unit_ids,
)A [row, bin, channel] representation and the reduction
that produced it.
A row is a unit where the calendar did the binning and a target where
a lookback did, since a unit carrying several targets cannot name a row
on its own. span and lag are set by a lookback
alone, and are what rebuilding one for new targets reads.
static names the channels holding the same number in every
bin, which flatten reads once each. position
names the channels calendar_channels made, which the
encoders read at their own amplitude rather than standardise.
bin_start, bin_end and
bin_partial are the calendar’s, and are None
on a representation the calendar did not bin.
coords is each row’s pair of coordinates and
unit_ids the identifier of the unit each row belongs to,
where timesift.timesift was given coords and
id. They place a row and are not channels, so no learner
reads them as a predictor, and a split of the rows splits them with
it.
Attributes:
-
values- np.ndarray -
units- tuple[str, …] -
bins- tuple[str, …] -
stats- tuple[str, …] -
grain- str -
year_start- str | None -
bin_start- np.ndarray | None -
bin_end- np.ndarray | None -
bin_n- np.ndarray -
bin_partial- np.ndarray | None -
span- int | None -
lag- int | None -
static- tuple[str, …] -
position- tuple[str, …] -
coords- np.ndarray | None -
unit_ids- tuple[str, …] | None