Python: combining the candidates
Source:vignettes/articles/python-combining.Rmd
python-combining.RmdWeights fitted on the out-of-fold predictions alone.
ensemble()
ensemble(
method: str = 'stack',
scope: str = 'all',
metric=None,
response: str | None = None,
min_score: float | None = None,
decay=None,
rule: str | None = None,
)Ask for an ensemble of the candidates a fit produced.
stack fits non-negative weights summing to one on the
out-of-fold predictions, mean and median
combine without fitting, and weighted uses each candidate’s
own mean score, its positive part rescaled to sum to one; where no
candidate scores above zero every candidate weighs the same.
decay, biomod2’s EMwmean.decay, changes that:
given a number d, the K candidates scoring
above zero take d**K for the best down to d**1
for the K-th, candidates on the same score share the mean
of their ranks’ weights, and a candidate at or below zero takes
none.
committee is biomod2’s committee averaging: each member
cuts each response at the threshold decision_threshold()
learns under rule from that member’s out-of-fold
predictions of every target, and the combination is the share of members
voting presence. A member holding no cut on a response does not vote on
it.
min_score, biomod2’s metric.select.thresh,
leaves a candidate whose mean score is below it ineligible, before
scope picks among what is left. scope is which
candidates are eligible: every one of them, only the several learners
sharing the best candidate’s representation, or only its learner across
the representations. metric names the metric the
eligibility, min_score and the weighted weights are read
by, or None for the scores the run already carries, and
response is the registered head whose loss the weights
minimise, or None for the head the run was fitted under.
Naming a head the run does not fit toward is an error rather than an
override, and ensemble_fit called on its own reads
None as "presence_absence".
ensemble_fit()
Fit the combiner on the out-of-fold predictions and nothing else.
oof is one [target, response] matrix per
candidate, in the response’s own row order. Only the cells the mask
admits are read, so every candidate is weighted on the same cells its
score was read on.
The weights are fitted to the response on those predictions, so the
combination scored against the same response is scored on the data its
weights were fitted to, and that score is optimistic.
timesift evaluates the stack the other way: each outer
fold’s weights are fitted on inner out-of-fold predictions of its
training targets and applied to the outer test fold.
ensemble_spread()
How far the members of an ensemble disagree: biomod2’s
EMcv and EMci.
Returns an [n, response, statistic] array, the
statistics being SPREAD_STATISTICS: the weighted mean
m of the members’ predictions under the stack’s weights;
their weighted standard deviation s, the square root of
sum(w (p - m)**2) / (1 - sum(w**2)), which is the sample
standard deviation when the weights are equal; the coefficient of
variation s / m; and the interval
m -+ t(1 - alpha / 2, n - 1) s sqrt(sum(w**2)),
n the number of members carrying weight, the t interval of
a mean of n members when the weights are equal. Under a
head whose predictions are probabilities the interval is held inside
zero and one. A committee’s and a median’s members are read at equal
weight, and a committee’s spread is that of the members’ predictions
rather than of their votes. sd, cv and the
interval are NaN where fewer than two members carry weight.
ensemble_weights()
The weight the combiner gave each of its members, or nothing where a run combined none.
EnsembleSpec
How the candidates are to be combined, which of them are eligible, and under which head.
Attributes:
-
method- str -
scope- str -
metric- object -
response- str | None -
min_score- float | None -
decay- object -
rule- str | None
Stack
A fitted combiner: what it does and what weight it gave each of its members.
A committee also carries each member’s cut on each response,
thresholds[member, variable] in the response’s own variable
order, NaN where a member holds no cut.
Attributes:
-
method- str -
weights- dict -
members- tuple[str, …] -
loss- str -
thresholds- np.ndarray | None -
variables- tuple[str, …] | None