Math Parser

Analysis › Process · version 1 · has a preview

A new localization field from an expression in the existing ones.

What it does

The Math Parser makes a new column of the localization table from an expression in the columns it already has:

on_time_ms = n_in_group * 20
z_true_nm  = z_nm * 0.8
within     = (xy_err_nm < 25) & (sigma_nm > 100)

The result is an ordinary column. The filter, the renderer, the colour coding and every other plugin use it like any column the fitter wrote. That is what makes the plugin useful for everything there is no button for: a refractive-index correction on z, an on-time in milliseconds, a ratio of two fitted quantities, or a flag (0 or 1) marking the localizations that a range filter on one column cannot express.

The expression is kept with the table, as a recipe. That matters once the localizations are grouped into blinks: the recipe says what the new column means for a blink, so it is not averaged by accident, and it is still there, with how it was computed, when the file is opened again.

How it works

1. Reading the expression. The text after = is read as a formula, not run as a program: only the column names, numbers, operators and functions of the table under In detail are allowed, and anything else is refused with a message pointing at the piece that is wrong. Every column the expression names must be in the table; if one is missing, the message lists the columns the table does have.

2. Computing it. The formula is evaluated over whole columns at once, one value per localization. A comparison gives 1 where it holds and 0 where it does not, so within above is a flag. A result that is a single number – a constant, or median(z_nm) – is copied to every localization.

3. Writing it. With apply to set to all localizations, the column is written for the whole table. With the selection, only the localizations of the current selection get the new value (the layer's filter, the ROI and the slab together); the others keep what the column had, or get no value (NaN) if the column is new.

4. The recipe. The expression is stored in the table together with grouped, the rule for what the column is on the grouped table. Running again into the same new field replaces its recipe. An expression that reads the field it writes – photons = photons * 2 – is applied once and kept as no recipe at all: re-evaluated at every grouping it would compound. The column then goes on being combined per blink by its own rule (photons are summed).

5. On the grouped table. When localizations are linked into blinks, each ordinary column is combined by a fixed rule (positions averaged, photons summed). A computed column instead follows its recipe:

The same column can mean quite different things per blink, which is why the choice is asked for rather than guessed:

The flag bright = photons > 1000 on a simulated table, taken to the grouped table by four rules. Left: the fraction of blinks that come out bright – recomputing asks whether the blink's summed photons exceed 1000, which more blinks do than have all their frames above 1000. Right: with mean, the flag becomes the fraction of a blink's frames that were bright, shown for the blinks of more than one frame.

In detail

What an expression may contain. Column names are the table's own (x_nm, photons, xy_err_nm, ...; the new field of an earlier run counts as one). Everything else:

what allowed
numbers20, 0.8, 1e3; constants pi, e, nan, inf, True, False
arithmetic+ - * / // (integer division) % (remainder) ** (power), unary - and +
comparisons< <= > >= == !=, one per bracket
combining flags& (and), ^ (exclusive or), ~ (not), and the vertical bar for or
choicea if condition else b, per localization
elementwise functionssqrt, abs, exp, log, log2, log10, sin, cos, tan, arcsin, arccos, arctan, arctan2, hypot, floor, ceil, round, sign, mod, rem, minimum, maximum, clip, where, isfinite, isnan, float, int
reductions (one number)mean, median, std, sum, min, max, percentile, count

Or is written with a vertical bar: (photons > 5000) | (sigma_nm > 150). The functions are numpy's (sqrt is numpy.sqrt, rem is numpy.remainder; float and int convert the type), called with their arguments in order: clip(photons, 0, 5000), percentile(z_nm, 90), where(z_nm > 0, 1, -1). The reductions ignore NaN, as numpy.nanmedian does; count is the number of values, NaN included. With apply to set to the selection, the expression is evaluated on the selected localizations only, so a reduction is the selection's own; recomputed on the grouped table, it is the grouped table's.

What is refused, and why. Each of these is a mistake that numpy would answer with either a crash far from the cause or, worse, a plausible wrong number, so it is caught with a message instead:

The expression is walked piece by piece rather than handed to Python's eval. It is saved with the table, in the history and in chain files, so it travels between people, and evaluating a string from someone else's file would let it do anything.

The result as a column. A boolean result is stored as 0/1 in float32, so the filter can take a range of it and grouping can reduce it. A floating-point result is stored as float32, like the rest of the table. An integer result – frame // 100 – keeps its integer type, since a frame number stops being exact in float32 above . A single number is copied to every row; a result of the wrong length, or one that is not a number, is refused.

The new field's name. It must be a word of letters, digits and underscores that does not start with a digit, so that a later expression can use it. The name of one of the functions is refused, and so are group_id and n_in_group, which grouping writes and would overwrite.

The recipe. Stored in the table's metadata, metadata["derived"], in the order the fields were defined (a field may be computed from an earlier one), as the field, the expression and the rule. It is saved in the file with the table. When the table is grouped, a computed column is left out of the usual per-column rules and either reduced by its rule – mean is the plain average of the localizations' values, not weighted by their precision as the positions are; any and all give 1 where the largest or smallest value is not zero – or recomputed from the expression on the grouped table. Each time the table is grouped, the expressions are also evaluated again on the ungrouped table, which is what keeps a field in n_in_group up to date with the latest linking. An expression in n_in_group or group_id needs the table to have been grouped once, since grouping writes those columns; before that the plugin says so. A recipe that cannot be recomputed because a column is missing is skipped with a message, and the table is still built.

A field computed for the selection is not a function of the table, so its expression is never recomputed. If grouped is recompute the expression, the rule used instead is mean, and the text says so.

The history. The last 20 pairs of field and expression are kept in math_parser.yaml in the settings directory, most recent first, each pair once, and offered under history. A history that cannot be read or written is treated as empty; it is never an error.

Parameters

settingdefaultwhat it does
new field
field
withinThe column the result is written to.

Choose a name that says what the number is, with its unit, like on_time_ms: it is what the filter and the colour coding will show.

=
expression
(xy_err_nm < 25) & (sigma_nm > 100)An expression in the column names; & and | combine comparisons, each in brackets.

Use Preview first: it computes the expression and reports the range and median of the result without writing anything. A precedence mistake gives numbers, not an error.

apply to
where
all localizationsThe selection is the layer's filter, the ROI and the slab together; the rest of the field is NaN.

A field written only for the selection has NaN elsewhere, and is only reduced, never recomputed, on the grouped table: over the localizations of each blink that have a value, and NaN for a blink that has none.

Choices: all localizations; the selection
grouped
grouped
recompute the expressionWhat the field means once the localizations are grouped: the expression again, or a rule to combine their values.

recompute for an expression in n_in_group or in anything that is a property of the blink; a combining rule for a measurement per localization; any or all for a flag.

Choices: recompute the expression; mean of the localizations; sum of the localizations; minimum; maximum; any localization (flag); all localizations (flag); leave it off
history
recall
An expression used before; choosing one fills the field, the expression and the grouped rule.

Picking an entry also restores the grouped rule it was used with.

Output

A result whose range is not what was expected – all zeros for a flag, or values a thousand times too large – is the sign of a mistake in the expression; Preview shows it before the table has it.

Differences from SMAP

Based on SMAP's Process/Modify/MathParser (Ries 2020). What differs:

References