Two tables to a scored comparison of representations, and the prediction that follows from it.
timesift()
timesift(
targets,
series=None,
*,
y,
x=None,
id=None,
time=None,
target_time=None,
static=None,
coords=None,
learners=None,
sift=None,
ensemble=True,
resampling=None,
n_inner=5,
choose: str = 'argmax',
response: str = 'presence_absence',
metric=None,
control=None,
keep_fits: bool = False,
seed: int = 1,
verbose: bool = True,
refit: bool = True,
)Compare every learner across every representation, and estimate choosing among them.
targets is one row per thing to predict and
series is the long, time-stamped record belonging to it;
both are mappings of column name to array, which a data frame satisfies.
y, x, static and
coords are selections over their own table: a name, a list
of names, a glob such as "sp_*", or a function of a name.
coords names the two columns holding each target’s
coordinates; they place a target and are not predictors, and a learner
that places a spatial field by them, timesift.hierarchical,
reads them off the array.
Within each outer fold of resampling the training
targets are split again into n_inner folds. Every candidate
is cross-validated on that inner split, choose picks one on
its inner score ("argmax" or
"coarsest_adequate", as in select_grain), and
the stack’s weights are fitted on the inner out-of-fold predictions.
Every candidate is then refitted on the whole outer training set and
predicts the outer test fold, and the selected candidate’s prediction
and the prediction combined under that fold’s weights are kept. Nothing
the outer test fold holds enters the choice or the weights it is scored
under, so estimate is of the procedure, selection and
stacking included. Its interval is across the response variables of this
dataset, all fitted and scored on the same targets and folds, and not
one for a new sample.
The same refits give every candidate an out-of-fold prediction on the
outer folds, which scores holds: the comparison, whose
highest level was picked out on the folds it is scored on.
n_inner=None runs no inner search and makes no estimate.
choice, models and stack are the
procedure applied to every target, for prediction: the rule read on the
outer scores and weights fitted on the outer out-of-fold
predictions.
Columns of targets that are neither the response nor the
identifier nor the anchor are ignored unless static names
them: a predictor is never picked up because it happened to be in the
table.
With repeats above one in resampling the
run is made once per repeat: scores carries a
repeat array and every repeat’s folds under numbers of
their own, estimate is read off all the repeats’ held-out
predictions, oof is a target’s mean prediction over the
repeats, and the stack is fitted on the repeats’ out-of-fold predictions
together. The models, the per-fold fits and folds are the
first repeat’s. refit says whether the candidates are
refitted on all targets at the end, which is what predict
reads; a repeated resampling asks it of its first run alone.
Timesift
Timesift(
candidates,
scores,
oof,
representations,
sift,
stack,
weights,
models,
folds,
cells,
y,
metric,
scorer,
response,
spec,
fits,
control,
estimate,
selected,
inner,
fold_weights,
predictions,
choice,
repeats,
)A fitted sift: every candidate’s out-of-fold predictions, their scores, and the combiner.
candidates and scores are tables held as
columns; oof, models and fits are
keyed by candidate name. What reads them – the summary, the weights, the
occlusion profile – reads them and never a model, so a number reported
here was read where the score was.
Attributes:
-
candidates- dict -
scores- dict -
oof- dict -
representations- dict -
sift- Sift -
stack- object -
weights- object -
models- dict -
folds- Folds -
cells- object -
y- Response -
metric- str -
scorer- object -
response- str -
spec- TimesiftSpec -
fits- dict -
control- object -
estimate- list | None -
selected- list | None -
inner- list | None -
fold_weights- list | None -
predictions- dict | None -
choice- str | None -
repeats- int
predict()
predict(
self,
targets,
series=None,
candidate: str = 'ensemble',
type: str = 'response',
rule: str = 'youden',
alpha: float = 0.05,
)Predict new targets, rebuilding each member’s representation from the stored settings.
Every candidate is refitted on all the targets at the end of a fit,
so what predicts here is one model per candidate rather than a fold’s
worth of them. "ensemble" combines them under the weights
fitted on every target, "selected" predicts with the
candidate the rule chose on every target (choice), and any
other value names one candidate.
type="binary" cuts each response into presence and
absence at the threshold decision_threshold() learns by
rule from the same candidate’s out-of-fold predictions of
the fit’s own targets, and returns 0.0 and 1.0, NaN for a response whose
held-out predictions give no cut. A binary map of a species is this
prediction on one target per map cell.
type="spread" reads the ensemble’s members side by side
rather than combining them: ensemble_spread() gives their
weighted mean, standard deviation, coefficient of variation and interval
at alpha as an [n, response, statistic] array,
biomod2’s EMcv and EMci, and on one target per
map cell an uncertainty map.
TimesiftSpec
How a fit was asked for: the columns each table plays, and the calendar they are read in.
Attributes:
-
y- tuple[str, …] -
x- tuple[str, …] -
id- str | None -
time- str | None -
target_time- str | None -
static- tuple[str, …] -
response- str -
metric- object -
coords- tuple[str, …]
candidate_table()
One row per candidate: its level, how many responses it was best on, and how it covered them.
A candidate whose learner and representation could not be paired carries no level, and is listed under the ones that do.
procedure_table()
The selected candidate’s and the stack’s held-out level under the run’s own metric, with the standard error across responses: one row each, or none where the run made no estimate.
plot()
Draw a ladder, a run or a selection, and return the table the drawing was made from.
A ladder is drawn as one line per learner across the grains, at the across-variable mean of the per-variable score, with a 95 percent interval from its standard error across variables on Student’s t with one degree of freedom fewer than there are variables; an open circle marks each learner’s best grain. A run is drawn the same way across the representations, with the stack’s held-out score across them as a dashed line: the curves are scored on the folds a choice among them would be judged on, so the best of them sits a little high, and the stack’s line does not. A selection is drawn as every candidate’s inner score in every outer fold, one line per fold, an open circle on the candidate the fold chose.
col is one colour per line, recycled. ax is
the matplotlib axes to draw on, a new figure’s where it is left unset,
and kwargs reach ax.set(), so
title= or ylim= are set as they would be
there. interval draws the interval on a ladder or a run and
is ignored on a selection, which draws none. Needs matplotlib.
project()
project(
fit,
series=None,
static=None,
candidate: str = 'ensemble',
type: str = 'response',
chunk: int = 5000,
**kwargs,
)Predict one target per cell of a grid, each carrying the record of its own cell.
This is BIOMOD_Projection() and
BIOMOD_EnsembleForecasting(): a map is the fit applied to
one target per cell, at the grain the fit reads. series is
an xarray.DataArray with a time dimension and
two spatial ones, or a mapping of the fit’s x column names
to such; static is an xarray.Dataset, or a
mapping of the fit’s static column names to
DataArray over the two spatial dimensions. Every input
shares one grid, and one set of instants. Cells are predicted in chunks
of chunk, and a cell is predicted where every input holds a
value at that cell: a cell with a missing reading anywhere is NaN in
every layer. candidate, type and the other
keywords are those of Timesift.predict.
Returns an xarray.DataArray over response
and the two spatial dimensions of the inputs; for
type="spread" a statistic dimension as
well.
range_change()
How the range of each response changes between two maps.
Counts the cells a response is lost from, kept in and gained, as
BIOMOD_RangeSize() does, a cell being in the range where
the response is predicted present. With L the cells lost,
K kept, G gained and A absent in
both: the current range is L + K, the later one
K + G, percent_loss is
100 L / (L + K), percent_gain is
100 G / (L + K) and change is
percent_gain - percent_loss. A cell that is NaN in either
map is left out of every count. now and later
are xarray.DataArray with a response
dimension, such as project() returns, or arrays of cells by
responses. threshold is None where the maps are already 0
and 1, and otherwise one cut per response, or one for all.
RangeChange
The change in range of each response: a table of counts,
and the map of codes (-2 lost, -1 kept, 0 absent in both, 1
gained) of the shape of the inputs.
Attributes:
-
table- list -
map- np.ndarray -
variables- tuple
pseudo_absences()
pseudo_absences(
pool,
presences,
n: int,
strategy: str = 'random',
id=None,
env=None,
quantile: float = 0.025,
coords=None,
dist_min: float = 0.0,
dist_max: float = np.inf,
lonlat: bool = False,
repeats: int = 1,
seed: int = 1,
)Draw pseudo-absences from a pool of background units.
Presence-only records give a response of ones. A model needs units
where the response is zero, and these are drawn from pool,
a mapping of column to array of background units that carry the same
columns as the targets. The strategies are
bm_PseudoAbsences()’s: "random" admits any
unit of the pool that is not a presence; "sre" a unit
outside the envelope of the presences, the band between the
quantile and 1 - quantile quantiles of each of
the env columns over the presences, a unit being outside
where it leaves the band in at least one column; "disk" a
unit whose distance to the nearest presence lies between
dist_min and dist_max, both included, on the
two coords columns, planar in the units of the coordinates
or in metres from longitude and latitude in degrees where
lonlat is true, on a sphere of radius 6371008.8 m.
A presence is never its own absence: a unit of the pool whose
id is one of the presences’ is not a candidate. The units a
strategy admits are the same in both languages; which n of
them are drawn depends on the language’s generator, as the folds of
fold_map do. The result is a mapping of column to array
holding the drawn rows of the pool, with set (the draw,
from 1) and pseudo (true) added; draw r uses
seed + r - 1.
select_columns()
The columns a selection names, in the order the selection names them.
An explicit list is taken in the caller’s order, because naming the response columns in an order is a decision; a glob and a predicate are taken in the table’s own order, because matching is not an ordering.
exclude names columns that exist but cannot be selected
here, so naming one is refused with the reason rather than reported as a
column that does not exist.
target_labels()
What names the rows of every array in one fit.
The unit identifier names them where a unit carries one target. Where
target_time lets a unit carry several, no identifier tells
them apart, so the row’s own position does, which is what
timesift.representation.lookback_matrix already names its
targets by.