Checks / ordering
Softmax / Sigmoid directly before Output
Check R07. Runs in the editor as you build, in CI through the
GitHub Action, and over the wire at
POST /api/v1/check. Milliseconds, before any GPU is billed.
info
ordering
R07
| Trigger | An explicit Softmax or Sigmoid layer is wired directly into Output. |
|---|---|
| Why | PyTorch's nn.CrossEntropyLoss applies LogSoftmax internally; an explicit Softmax before it double-applies and slows training. BCEWithLogitsLoss has the same issue with Sigmoid. |
| Source | torch.nn.CrossEntropyLoss docs. |
Why it is not a lint you can ignore
A structural mistake does not fail at review time and it does not fail at import
time. It fails when the module is constructed on the training node, after the job was queued and
the dataset was downloaded. That is why this runs before the spend and not after it.
Run this check on your own model
Free, no account needed
Every check
41 structural checks: 6 guardrail gates and 35 architecture advisor rules. See the full catalogue.
← R06 Dropout directly before BatchNorm ยท R08 Normalization at output →