Expressions#
Every expression: in the file — a constraint's, the objective's, a named
quantity's — is written in one small arithmetic language:
expression ::= arithmetic | arithmetic COMPARATOR arithmetic
arithmetic ::= atom | unary_op arithmetic | arithmetic binary_op arithmetic
| function_call | "(" arithmetic ")"
atom ::= NUMBER | NAME
unary_op ::= "+" | "-" binary_op ::= "+" | "-" | "*" | "/" | "**"
COMPARATOR ::= "<=" | ">=" | "=="
function_call ::= NAME "(" [pos_arg ("," pos_arg)*] ["," kwarg ("," kwarg)*] ")"
kwarg ::= NAME "=" (arithmetic | QUOTED | "[" NAME ("," NAME)* "]")
NAME ::= [a-zA-Z][a-zA-Z0-9_]*
NUMBER ::= integer | float | "inf" | ".inf"
Precedence, highest first: **, then unary + -, then * /, then binary
+ - — so -x ** 2 is -(x ** 2) and -x * y is (-x) * y, as in Python.
Parentheses override. A float may carry an exponent (1e5, 2.5e-3); a sign
is always the unary operator's. A keyword given twice in one call is an error
rather than the later one winning.
Degree 2 in the math, degree 1 beside it#
The objective and constraints take variable * variable. So a quadratic
cost, sum(p * p * wear, over=g), and a quadratic row, p * q >= floor, are
both sayable and both say what they mean.
Three rules bound it:
- At most one factor may be a sum of terms.
sum(p, over=g) * sum(q, over=g)is refused: that is every term of one against every term of the other, and nothing in the file says how many that is. Multiply before reducing, or give the reduction a name — a variable constrained to equal it is one term. Factors carrying different dims are fine:x * y * linkbroadcasts and joins through the table that couples them. - Degree stops at 2.
p * p * pis refused wherep * pis not. - Everything beside the math stays affine — a bound and a
piecewise:link. A bound is a number per column; a link expands into declarations that must themselves be affine. A named expression, by contrast, is read at the ceiling of wherever the math reads it — degree 2 in the objective or a constraint, affine in a piecewise link. One nothing in the math reads is held to no degree at all: it is a reported quantity (below), which is what lets it divide by a variable, cube one, or calldual().
/ needs a variable-free divisor everywhere, and a single factor rather than a
sum — both decided at load time, since neither depends on the numbers that
arrive, and a variable divisor is rational rather than polynomial, which is
outside the language at any degree.
** takes a base and an exponent that carry no variable, and nothing else.
growth ** period is a discount factor — one number per coordinate, folded
from a rate the model binds and a period it declares — so it is the arithmetic
* already does, spelled the way the maths is written. Two refusals bound it,
both at load:
- A variable anywhere under it.
x * xis how a square gets written; above degree 2 there is no rewrite at all. A variable exponent is out for a sharper reason —p ** nis affine atn = 1and quadratic atn = 2, so the degree would be a property of the data andto_speccould not answer with nothing bound. - An operand that adds. Addition does not distribute over
**, so(1 + rate) ** periodis two factors wearing one and is refused wheregrowth ** periodis not. Bind the factor itself.
What it costs is a consumer's question#
Saying it is one question; solving it is another (the ceiling). This language admits degree 2 in the objective and constraints and says nothing about which solver, lane or file format takes it — that is the consumer's axis, and each consumer answers for itself. Two things no consumer can answer from the model alone, because both are properties of the data: whether a quadratic form is convex, and whether a quadratic row can be priced.
A piecewise: block with method: convex remains the way to spend a curve and
keep the LP, its duals and its warm start.
Name resolution#
A name is a letter or an underscore, then letters, digits or underscores — the spelling an expression uses to refer to one. A declaration keyed by anything else is a load error, because nothing in the file could ever write it.
One flat namespace covers dimensions, parameters, variables, named
expressions, macros and the built-in operators. A collision is a load error
naming both declarations — there is no shadowing, because under it declaring a
parameter named snapshot would silently change what an existing
where: "snapshot > 0" means.
Position decides which kinds of name are legal, and every name's kind is fixed when the file loads:
| Position | Legal kinds |
|---|---|
expression (p * cost) |
variable, parameter — the parameter a number (dtype) |
dimension argument (over=) |
dimension |
lookup argument (by= on sum / at) |
lookup — never a dimension |
where string |
parameter, variable, dimension, lookup (where strings) |
bounds.lower / bounds.upper |
parameter name, or a number |
the edge key of shift |
'wrap' quoted, or a bare number; never a dimension |
dual argument (dual(c)) |
constraint — resolved against constraints alone, never the flat namespace (reported) |
A bare word in a keyword-argument value is a name to resolve, which is why
wrap is quoted: shift(x, over=wrap, edge='wrap') reads unambiguously even
where a dimension is called wrap. edge is the one keyword whose key is
fixed rather than naming a dimension, so a dimension called edge does not
change what it means.
A dimension in a value position is an error — it is a coordinate space, not data. To use its coordinates as data, declare a parameter over it.
A str or bool parameter there is an error too — data, but not a number.
A label selects and a flag masks, which is what a where is for; multiplying by
either is a cast the file never wrote, so only dtype: float and dtype: int
stand as a coefficient, a term or a divisor
(dtype).
Constraints are outside the flat namespace. The one position that names a
constraint is dual's argument,
resolved against constraints alone — so a bare name never reaches a constraint
and a model may still name a constraint after a variable. What reads a solve
back keys on the label space as well as the name for that reason. The objective
carries no name at all.
Dim algebra#
Parameter dims and variable foreach are declared, and dimension arguments
are name-checked, so every expression's dim set is known before any data
binds:
| Node | Dim set | Error |
|---|---|---|
| number | {} |
|
| parameter / variable | its dims / its foreach |
|
-x, +x |
dims(x) |
|
a + b, a * b, a / b |
dims(a) ∪ dims(b) |
|
sum(x) |
{} |
if dims(x) is empty already |
sum(x, over=d) |
dims(x) − {d} |
if d ∉ dims(x) |
sum(x, by=l) |
(dims(x) − {over(l)}) ∪ {into(l)} |
if over(l) ∉ dims(x), or into(l) ∈ dims(x) already |
sum(x, by=[l, m]) |
(dims(x) − {over(l)}) ∪ {into(l), into(m)} |
the same, plus: if l and m are over different dims, or target the same one |
at(x, by=l) |
(dims(x) − {into(l)}) ∪ {over(l)} |
if into(l) ∉ dims(x), or over(l) ∈ dims(x) already |
shift(x, over=d, offset=n) |
dims(x) |
if d ∉ dims(x) |
Binary operators union: an outer product is legitimate when the frame declares the result. What must not be silent is a declaration that disagrees, so:
- a constraint requires
dims(lhs) ∪ dims(rhs)to equal itsforeach. A stray dim multiplies rows and an unusedforeachdim repeats one row across them — either way you would build a different model than the file reads as; - an objective must carry no dims at all — it is one number, and the sums that make it one are written in the expression;
- a
wherepredicate's dims and a bound parameter's dims must not exceed the frame they sit in.
Get it wrong and you are told at load time, not at solve time.
Where strings#
A where: is a boolean mask, and true means "this coordinate exists".
where_expr ::= atom | "NOT" where_expr | where_expr ("AND"|"OR") where_expr
| "(" where_expr ")"
atom ::= NAME | NAME COMPARATOR value | POSITION COMPARATOR INTEGER
| "True" | "False"
COMPARATOR ::= "<=" | ">=" | "==" | "!=" | "<" | ">"
value ::= NUMBER | QUOTED | NAME_OR_STRING
POSITION ::= "position" "(" NAME [ "," "by" "=" NAME ] ")"
QUOTED ::= "'" chars "'" | '"' chars '"'
| Surface | Names a… | Meaning |
|---|---|---|
name (bare) |
parameter | what defined means is the declaration's to say: a bool is its own answer, a str is defined wherever the table has a row, and a number has to be finite as well — 0.0 counts, inf does not, though it is a value everywhere else |
name (bare) |
variable | the variable exists at this coordinate — the counterpart of the parameter row, and how you say which coordinates the row-dropping rule applies to |
name (bare) |
dimension | load error: it is true everywhere, so it reads as a condition and is not one. Compare it instead |
name OP value |
parameter | element-wise; a null compares false. The right-hand side is a literal number, or a bare name read as a string coordinate |
name OP value |
dimension | a filter on the frame's own coordinate column |
name (bare) |
lookup | defined: the label maps somewhere. A lookup may be partial, and this is how a declaration asks for the labels that do map |
name OP value |
lookup | a filter on the lookup's column of its over dimension's index — which therefore has to be in the frame. A null value is false, whatever the comparator |
name OP name |
two lookups | the one comparison whose both sides are structure. Legal only where both map out of the same dimension and into the same one — from != to excludes a self-loop |
position(name) OP i |
one dimension | where the row sits along that dimension's own order, as an integer — 0 is first, negative counts from the end. Both sides are integers, so every comparator reads the one way |
position(name, by=lookup) OP i |
a dimension and a lookup over it | the same, counted within each group the lookup makes — every period's first snapshot, whatever each period's length |
AND OR NOT |
— | case-insensitive; NOT binds tighter than AND, which binds tighter than OR |
True / False |
— | literals, decided at load wherever they stand: True is the same as no where, False is a declaration with no rows, and one under an AND or an OR settles that side — x AND False is the declaration with no rows too. A double negation goes the same way, NOT NOT x being x, so what a page prints is what the mask decides rather than how it was spelled. A case when: is the one place a mask that folds to a literal is refused instead |
The mask's dims must not exceed the frame it sits in (dim algebra), and an undeclared bare name is a load error.
Defined is not non-zero, and the difference is a property of the data rather
than of the model. A bare parameter name is true wherever the table has a row,
0.0 included — so one where: masks nothing against a table padded with zeros
and deletes rows against a sparse one carrying the same information. Where the
intent is non-zero, compare for it: where: "inflow != 0" rather than
where: inflow, which a padded zero satisfies.
Comparing two parameters is not in the language — precompute a boolean parameter in data prep — and neither is comparing two dimensions. Two lookups are the exception, and only two that share both ends: over one dimension they are two columns of one index, so the comparison is a filter on that table rather than a join between two, and into one dimension they draw from one label set, so a match is possible at all. Over different dimensions no row carries both, and into different label sets no value can ever match — either is a load error. A label space owns its values and is therefore never the other side of one.
The string reading of a right-hand-side name is for names the model does not declare, which is how a string coordinate is compared; a declared name there is a load error naming the near miss, because reading it as text would compare a coordinate column against another declaration's name and mask everything out.
Quote a label that is not an identifier, and quote a date. A bare word has
to look like a name, so combined-cycle, IT-north and CCGT 400MW are only
sayable in quotes — and quoting is also what says label, not name, so a
quoted word is never read as a declaration and never a near-miss error.
A comparison is checked against the declared dtype. This matters most for
dates: a datetime dimension compared to a number is compared against the
epoch, so snapshot > 0 would silently mean "after 1970-01-01". That is a
load error naming the fix. A datetime boundary is a quoted ISO date —
snapshot > '2030-01-01', or '2030-01-01T06:00' with a time. Calendar
arithmetic, resampling and timezone conversion stay data prep.
position(dim) converts a dimension to where the row sits along it, so a
boundary clause survives the index being relabelled:
dimensions:
snapshot: { dtype: int }
parameters:
soc_initial: { dims: [] }
variables:
soc: { foreach: [snapshot], bounds: { lower: 0 } }
constraints:
soc_start:
foreach: [snapshot]
where: "position(snapshot) == 0" # not: snapshot == 0
expression: soc == soc_initial
A recurrence needs its first position seeded, and the label that happens to be
there is a property of the data — relabel [0, 1, 2] to [1, 2, 3] and
snapshot == 0 matches nothing, leaving the recurrence unanchored. -1 is the
last position, -2 the one before it. A position no coordinate occupies is
an error at bind, not an empty mask: the clause exists to seed a row, and
seeding none is the failure it was written to prevent.
The order counted along is the dimension's own — the one shift walks, and the
one the index declares — not the bytewise order a label comparison uses.
The conversion is on the left, and that is what makes an ordering readable.
position(snapshot) > 0 is "not the first row", on any axis, because both
sides are integers. Naming the coordinate at a position and comparing
coordinates against it would have made the same clause mean either that or "a
coordinate sorting after the first one" — two different masks wherever the
coordinates do not arrive sorted, and nothing in a file says they do
(#32). A comparison of
values is still written against the dimension itself, where it always was:
snapshot > '2030-01-01'.
by= counts inside each group a lookup makes, which is the boundary a
multi-period model wants — one seeded row per period rather than one per
horizon:
dimensions:
snapshot: { dtype: int }
period: { dtype: int }
lookups:
period_of: { over: snapshot, into: period }
parameters:
soc_initial: { dims: [period] }
variables:
soc: { foreach: [snapshot], bounds: { lower: 0 } }
constraints:
soc_start:
foreach: [snapshot]
where: "position(snapshot, by=period_of) == 0"
expression: soc == at(soc_initial, by=period_of)
by= takes a lookup over the dimension being counted — groups a row of
that dimension is actually in. Unlike sum(by=) and at(by=)
it need not be a groupable one: counting inside a group lands no terms
anywhere, so a label space partitions the rows just as well (#280). A row reads its own group's boundary, the broadcast at(by=)
already defines, and -1 is each group's last however long that group is.
Periods of different lengths therefore need nothing special, which is the case
no single position along the whole axis can express.
A coordinate the lookup sends nowhere is in no group, so it is no group's boundary — the same reading a null value gets everywhere else. A group shorter than the position is an error at bind, for the reason the ungrouped form has one: a boundary naming no coordinate leaves those rows unseeded.
String labels order bytewise, whatever order the dimension declared them
in. Declaration order is a different axis — it is what shift walks — and a
where never reads it: node >= 'b' means the same thing however the nodes
were listed. A label the dimension does not carry compares equal to nothing, so
the mask is false there rather than an error: quoting already said label, not
name, and a label is data.
Named expressions#
A quantity the model names once and can read back after a solve:
dimensions:
generator: { dtype: str }
parameters:
rate: { dims: [generator] }
variables:
p: { foreach: [generator] }
expressions:
total_generation: sum(p, over=generator)
emissions:
expression: sum(p * rate, over=generator)
description: CO2 released, the quantity a cap would bound
Written as a bare string until it carries a description:, which is when it
gains the mapping form.
A named expression has fixed dims — they fall out of its body, so there is
no foreach — and an observable identity: after a solve,
a consumer can read its value back over its own dims.
That is the point of naming a
quantity: the CO₂ a constraint bounds and the CO₂ a summary reports are one
definition, validated once.
Where a constraint or the objective references one, it is substituted before anything consumes the model, so a reference costs nothing at build time. It is lowered only when it is read, so a model with fifty named expressions that reads none pays for none.
A named expression is one of two things, and the file never says which —
the objective and the constraints do. One they inline, directly or through
another entry or a macro, is in the math: it stands inside the program a
solver sees, held to the same
degree-2 ceiling the math holds to
everywhere else, where it is read. One nothing in the math names is
reported: a statistic the solver never sees, arithmetic over numbers a
solve has already produced, where those restrictions lift — which is what lets
it divide by a variable, cube one, or call
dual(). See
Reported expressions for what lifts, and how the split is
decided.
cases: — one quantity, a value per region#
Some quantities have no single expression. The commitment state a unit carries
into a snapshot has three regimes: 1 for a unit that is never switched off, an
initial condition at the first snapshot, and the last snapshot's status
everywhere else. Written at the constraint, those regimes fork the inequality
three ways. Named here, the inequality is written once:
expressions:
previous_status:
description: the commitment state a unit carries into a snapshot
foreach: [snapshot, generator]
cases:
always_on:
when: "not committable"
expression: 1
boundary:
when: "committable and position(snapshot) == 0"
expression: status_initial
otherwise: shift(status, over=snapshot, offset=1)
constraints:
ramp_up:
foreach: [snapshot, generator]
expression: >-
p - shift(p, over=snapshot, offset=1, edge=0)
<= ramp_limit * previous_status + start_up_limit * (1 - previous_status)
One case is one row, and otherwise: is the last:
The shape. A named expression carries exactly one of two things: an
expression:, or a cases: block. A cases: block is a map of named cases,
each with a when: and an expression:. Two keys sit beside it: the
otherwise:, which carries whatever the cases leave, and the foreach:, which
declares the dimensions all of them range over.
Those dimensions are the block's frame, and one point of it — one snapshot for one generator, in the example above — is a coordinate. Every rule below is about which case owns which coordinate.
The rules#
No two cases may claim one coordinate. One generator at one snapshot
cannot have two previous statuses, so a file where two when: masks can hold
at once is refused at load. The refusal comes before any data binds, and it
names the pair, a coordinate they both claim, and the rewrite:
Named expression 'previous_status': casesalways_onandboundaryboth claim the value where committable is false, the position of snapshot is 0. A coordinate two cases claim has two values, so it has none — narrow one of the twowhen:strings by the negation of the other, or drop the wider one and letotherwise:carry that region.
That is why boundary above says committable and.
A when: the data cannot decide is not a case. A mask the connectives
settle on their own — True, False, or anything that folds to one, like
committable OR True — states no condition for the data to answer, so it names
no region. Both halves are refused at load, and the refusal names the rewrite:
Named expression 'previous_status', case 'always_on': the mask admits every row, so no other arm can hold anywhere andotherwise:covers nothing. Write the expression withoutcases:, or narrow thewhen.
An always-false arm is the other half — it never applies, so delete it or widen
it. A declaration's where: is not held to this rule and cannot be: there
False is how a file says the declaration has no rows, and True is the same
as writing no mask at all. It is the when: on an arm that has to be a
question, because the arms are kept apart by proof.
A pair the check cannot decide is refused too, and that refusal names its
rewrite as well. The one that comes up is position(snapshot) == 0 against
position(snapshot) == -1. On an axis with a single member those two pick the
same row, and how many members an axis has is data rather than declaration. So
count from one end only.
otherwise: is the value wherever no when holds, and it takes every
coordinate the cases leave. It carries no mask of its own, so nothing narrows
the frame it is written against. It is the one value that has to hold up at
every coordinate — those where a parameter is absent or a label is unnamed
included.
Covering a coordinate is not the same as having a value there, and a case
that claims a coordinate may still be empty at it. The otherwise: above shows
how. Its shift carries no edge=, so it produces nothing at the first
snapshot; previous_status is whole there only because every unit at that
snapshot is claimed by boundary or by always_on instead. Close such a hole
in one of three ways:
- widen a
whenuntil it covers the coordinate the case drops out at, - give the
shiftanedge=, - or set
absence: zeroon a masked variable.
Nothing catches one left open at load, because whether a case has a value there depends on the data.
foreach: is required with cases and refused without. The dims of an
uncased expression fall out of its body. The dims of a cased one cannot,
because a case may be a single number while the condition that selects it
ranges over dimensions — always_on above is exactly that. So the frame is
declared. Each when: is held to it, the way a variable's or a constraint's
mask is, and each case's value must sit inside it.
The dims of a reference are the declared foreach, not the union of the
cases: one narrower than the frame broadcasts, exactly as a parameter with
fewer dims does.
cases: inside a macros: template is not supported. The otherwise:
would have to cover a frame the macro does not have until it is called.
Why it is shaped this way#
A coordinate with two values has no single value, so it is no longer a quantity. The regimes are therefore kept apart by proof rather than ranked by position. That is what makes the cases readable in any order: each one says where it applies on its own terms, without the ones above it in mind. A tool that re-sorts the keys of a file cannot change what the file means.
otherwise: is required because it makes the quantity whole without a second
proof: it carries no condition, so there is no condition on it to fail. A
coordinate that no when matched would have no value at all, and absence
spreads, so any constraint reading the expression would lose rows
it never masked.
It is written beside cases: rather than inside them because it is not a
region like they are. It is what is left over. And since it carries a value and
nothing else, it is written as a bare value — the same shorthand expressions:
itself takes.
How a cased expression prints#
Every other named expression is substituted where it is used, and prints nothing under its own name. A cased one is the exception, because it cannot be substituted legibly. Three cases are three rows tall, so whatever follows the name in the constraint would sit beside the middle row. The quantity would also print once per use, though the file writes it once.
So a use prints the symbol, and the block above prints once under a
Definitions heading between Subject to and Variable domains, in
declaration order, where a paper states a quantity defined by region.
A cased expression joins the symbol pool like any other quantity, so
--symbols can rename one. Uncased ones stay out, since a table entry for one
would never apply.
The unit commitment example is the whole model this section is drawn from.
Macros#
A parameterised template. It has no dims until it is called, and each call site may give it different ones — so it has no value a solve could report, and is never readable:
weighted_sum:
args: [array, weights] # positional formals, default []
kwargs: [over] # keyword formals, default []
template: sum(array * weights, over=over)
Both blocks hold arithmetic and no comparison. Arguments expand before substitution (call-by-value), so they may themselves use macros and named expressions. Formals shadow model names inside a template but may not collide with a declared dimension. Arity is checked per call site, and cycles are reported with the reference chain.
Every template is parsed and name-checked at load time even if it is never called — a macro nobody uses cannot hide a typo.
Anything composable out of the built-in operators belongs here. Math that is not sayable at all is out of scope (limits).