tests/test.sh exit code (0 → resolved); the classification below is post-hoc and cannot change it.List app structure
ls -la && echo "---SUBMISSION---" && ls -la submission/ && echo "---DATA---" && ls -la data/
total 89 drwxr-xr-x 4 user user 4096 Jun 30 22:55 . drwxr-xr-x 21 root root 4096 Jun 30 22:55 .. -rw-r--r-- 1 user user 14914 Jun 30 22:55 PROMPT.txt -rw-r--r-- 1 user user 5 Jun 30 22:55 agent.pid -rw-r--r-- 1 user user 0 Jun 30 22:55 agent.stderr -rw-r--r-- 1 user user 4498 Jun 30 22:56 agent_trajectory.jsonl drwxr-xr-x 2 user user 60 Jun 30 22:55 data -rw-r--r-- 1 user user 0 Jun 30 22:55 run.log -rwxr-xr-x 1 user user 45855 Jun 30 11:01 run_checks.py drwxr-xr-x 3 user user 128 Jun 30 22:55 submission -rwxr-xr-x 1 user user 1318 Jun 30 11:01 test.sh -rwxr-xr-x 1 user user 629 Jun 30 11:01 timer.sh ---SUBMISSION--- total 100 drwxr-xr-x 3 user user 128 Jun 30 22:55 . drwxr-xr-x 4 user user 4096 Jun 30 22:55 .. drwxr-xr-x 2 user user 4096 Jun 30 22:55 __pycache__ -rw-r--r-- 1 user user 48121 Jun 30 11:01 autograd.py -rw-r--r-- 1 user user 1692 Jun 30 11:01 data.py -rw-r--r-- 1 user user 23705 Jun 30 11:01 nn.py -rw-r--r-- 1 user user 12275 Jun 30 11:01 optim.py -rw-r--r-- 1 user user 2900 Jun 30 11:01 train.py ---DATA--- total 11108 drwxr-xr-x 2 user user 60 Jun 30 22:55 . drwxr-xr-x 4 user user 4096 Jun 30 22:55 .. -rw-r--r-- 1 user user 3953364 Jun 30 11:01 reweight.npz -rw-r--r-- 1 user user 7411924 Jun 30 11:01 train.npz
Count lines in files
wc -l submission/*.py run_checks.py test.sh 2>/dev/null
1306 submission/autograd.py
44 submission/data.py
543 submission/nn.py
272 submission/optim.py
65 submission/train.py
937 run_checks.py
24 test.sh
3191 total/app/submission/autograd.py
1 """Reverse-mode automatic differentiation over numpy arrays. 2 3 YOU MUST IMPLEMENT EVERYTHING MARKED ``# TODO`` / ``raise NotImplementedError`` using numpy 4 ONLY. No torch/jax/tensorflow/autograd/etc. The grader's conformance suite differentiates each 5 op below with central finite differences and compares against YOUR analytic backward, so every 6 backward must be correct (a single wrong reduction / Jacobian / group-weighting fails the checks). 7 8 A ``Tensor`` wraps one ``numpy.ndarray`` (stored as ``float64`` in ``self.data``) and records 9 the op that produced it so ``.backward()`` can accumulate ``.grad`` into every leaf. 10 11 CONTRACT (the grader builds its own reference against this exact API): 12 * ``Tensor(data, requires_grad=False)`` -- ``self.data`` is ``float64``; ``self.grad`` is None 13 until backward, then a same-shape ``float64`` array. requires_grad leaves accumulate grad. 14 * Every op returns a NEW Tensor whose ``requires_grad`` is the OR of its differentiable inputs, 15 and whose ``_backward`` closure does ``input._accum(local_grad)`` (so multiple paths add up). 16 * ``t.backward(grad=None)`` -- seeds ``grad`` (all-ones for a scalar) and walks the graph in 17 reverse topological order, calling each node's ``_backward``. Repeated backward accumulates. 18 * BROADCASTING: binary ops broadcast like numpy; the backward MUST reduce (sum) the upstream 19 gradient back to each input's ORIGINAL shape (use the provided ``_unbroadcast`` helper). 20 * …[truncated 52049 chars]
/app/submission/nn.py
1 """Neural-network layers + models + the robustness machinery, built on YOUR autograd engine. 2 3 Implement every ``# TODO`` forward (and the robustness-helper bodies). The parameters + 4 ``named_params`` naming + the checkpoint round-trip are already wired; you compose the autograd 5 ops. The grader checks each layer's forward against its OWN reference AND finite-difference-checks 6 the gradients that flow through your autograd, so the composition must be exactly right. 7 8 THE FINAL-METRIC MODEL is the pinned ``SmallCNN`` (a conv FEATURE EXTRACTOR + a single linear 9 CLASSIFIER). The split is what makes last-layer retraining (DFR) on a group-balanced set 10 expressible against the pinned graph: freeze the extractor and re-fit ONLY ``classifier``. 11 ``featurize(X, training)`` returns the post-ReLU embedding so you can freeze the extractor. 12 13 The other layers/models (Linear / Conv2d / BatchNorm2d / BatchNorm1d / LayerNorm / Dropout / 14 AvgPool2d / MLP / ResidualMLP / a TinyConvNet) are exercised by the conformance suite across a 15 WIDE surface -- each is an independent potential bug. The robustness machinery (``GroupDROState`` 16 exponentiated-gradient adversary-weight update, ``balanced_group_weights``, 17 ``class_balanced_weights``, ``log_class_prior``) is the domain-specific heart this task adds; the 18 ``robust`` milestone checks them. 19 20 SmallCNN naming contract (do not change -- the verifier indexes by these exact keys): 21 conv1.weight (Wd,3,3,3) | con …[truncated 24273 chars]
/app/submission/optim.py
1 """Optimizers + LR schedulers + grad clipping over a flat list of autograd ``Tensor`` parameters. 2 3 Implement the ``# TODO`` update math. All optimizers operate IN PLACE on ``param.data`` using 4 ``param.grad``. ``zero_grad()`` resets every grad to None. ``t`` is the 1-based step count. The 5 grader runs one (and several) steps of each and compares your updated params against its 6 reference within a tight tolerance, so the math must be EXACTLY right -- every optimizer / 7 scheduler / clipper is an independent potential bug. 8 9 The training recipe uses AdamW (decoupled weight decay) + ``clip_grad_norm`` for the ERM 10 extractor and the last-layer (DFR) retrain; the rest of the zoo is graded for parity. 11 12 SGD : g = grad + wd*p; v = mu*v + g; upd = (g + mu*v) if nesterov else v; p -= lr*upd 13 Adagrad : g = grad + wd*p; s += g^2; p -= lr*g/(sqrt(s)+eps) 14 RMSprop : g = grad + wd*p; s = alpha*s + (1-alpha)*g^2; p -= lr*g/(sqrt(s)+eps) 15 Adam : g = grad + wd*p (COUPLED L2); m,v EMA; mhat=m/(1-b1^t); vhat=v/(1-b2^t); 16 p -= lr*mhat/(sqrt(vhat)+eps) 17 AdamW : DECOUPLED wd; m,v EMA; p -= lr*( mhat/(sqrt(vhat)+eps) + wd*p ) 18 Adamax : m=b1*m+(1-b1)*g; u=max(b2*u, |g|); p -= (lr/(1-b1^t))*m/(u+eps) (infinity-norm) 19 NAdam : m,v EMA; mhat=m/(1-b1^t); vhat=v/(1-b2^t); 20 mnes = b1*mhat + (1-b1)*g/(1-b1^t); p -= lr*mnes/(sqrt(vhat)+eps) 21 Lion : upd = sign(b1*m + (1-b1)*g); p -= lr*(upd + wd*p); m = b2*m + (1-b2 …[truncated 11759 chars]
/app/submission/data.py
1 """Data loading + group helpers for the worst-group / spurious-correlation task.
2
3 Implement the ``# TODO`` bodies (numpy only). The splits ship as ``.npz`` files with arrays
4 ``X`` (N,3,S,S) float, ``y`` (the SHAPE class to predict) and ``g`` (the spurious background-color
5 id). A "group" is the pair ``(y, g)``; ``group_ids(y, g, n_classes) = y*n_classes + g`` flattens it
6 to a single id in ``[0, n_classes**2)``.
7 """
8 from __future__ import annotations
9
10 from typing import Dict, Tuple
11
12 import numpy as np
13
14
15 def load_npz(path: str) -> Tuple[np.ndarray, np.ndarray, np.ndarray]:
16 """Load (X, y, g) from a .npz; return float64 X, int64 y, int64 g. (Provided.)"""
17 d = np.load(path)
18 return (np.asarray(d["X"], dtype=np.float64),
19 np.asarray(d["y"], dtype=np.int64),
20 np.asarray(d["g"], dtype=np.int64))
21
22
23 def group_ids(y: np.ndarray, g: np.ndarray, n_classes: int) -> np.ndarray:
24 """Flatten (y, g) to a single group id y*n_classes + g (int64)."""
25 # TODO
26 raise NotImplementedError("group_ids")
27
28
29 def iterate(X, y, g, batch_size, rng):
30 """Yield (Xb, yb, gb) minibatches over a random permutation of the dataset."""
31 # TODO
32 raise NotImplementedError("iterate")
33
34
35 def worst_group_acc(preds, y, g, n_classes) -> Tuple[float, float, Dict[int, float]]:
36 """Return (min over (y,g) groups of within-group accuracy, average accuracy, per-group dict)."""
37
…[truncated 318 chars]/app/submission/train.py
1 """Training entry point. Implement your worst-group-robust recipe, train the pinned SmallCNN,
2 and save the checkpoint the grader will load.
3
4 Run it however you like (e.g. ``python3 /app/submission/train.py``). It MUST end by writing
5 ``/app/submission/model.npz`` via ``save_checkpoint`` below. The grader loads that file into its
6 OWN reference SmallCNN (architecture pinned to WIDTH/EMBED_DIM below -- keep these) and evaluates
7 WORST-GROUP accuracy on a SEALED, group-balanced test set; you pass the metric part iff
8 worst-group accuracy >= the frozen threshold. (You ALSO must pass the component conformance suite
9 over autograd.py / nn.py / optim.py.)
10
11 A naive cross-entropy ("ERM") run on the correlated train split will reach high AVERAGE accuracy
12 but ~0 worst-group accuracy -- it keys on the spurious background color. Removing that reliance is
13 the task. One well-known cheap recipe: train the feature extractor on the correlated data, then
14 RE-FIT ONLY the final ``classifier`` on the GROUP-BALANCED reweight split with a group-balanced
15 (group-weighted) loss so the head stops keying off color. You are free to use any method, as long
16 as you submit a state_dict for the pinned architecture.
17 """
18 from __future__ import annotations
19
20 import argparse
21 import json
22 import os
23 import sys
24
25 import numpy as np
26
27 sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
28 sys.path.insert(0, "/app")
29 sys.path.insert(0, "/app/su
…[truncated 1589 chars]/app/run_checks.py
1 #!/usr/bin/env python3 2 """PUBLIC component check runner for the from-scratch numpy framework (worst-group / spurious- 3 correlation task). 4 5 Run this AS OFTEN AS YOU LIKE while you implement to get a per-check progress signal: 6 7 python3 /app/run_checks.py # check the framework in /app/submission 8 python3 /app/run_checks.py --submission /dir 9 python3 /app/run_checks.py --only robust # focus a single milestone 10 11 It imports YOUR modules from /app/submission (autograd.py / nn.py / optim.py / data.py), exactly 12 as the grader does, and runs a PUBLIC portion of the component conformance suite, printing a clear 13 per-check report. Each FAIL line names the exact op / layer / optimizer / robustness-loss that is 14 wrong. 15 16 HOW THIS RELATES TO YOUR GRADE 17 ------------------------------ 18 This runner is a DEBUGGING AID, not the grader. It checks your framework against self-consistent 19 references it computes on its own: 20 * autograd_grads : central FINITE-DIFFERENCE of YOUR forward vs YOUR analytic backward. 21 * nn_forward : mathematical PROPERTIES your forward must satisfy. 22 * optim : YOUR optimizer step vs the closed-form update re-derived here in numpy, for 23 every optimizer + the LR schedulers + both grad clippers. 24 * robust : finite-diff of YOUR robustness-loss backward + the per-group-normalization / 25 Group-DRO / weight-helper PROPERTI …[truncated 47999 chars]
Inspect data
python3 -c "
import numpy as np
d=np.load('data/train.npz'); print('train', {k:(d[k].shape,d[k].dtype) for k in d})
print('y vals', np.unique(d['y']), 'g vals', np.unique(d['g']))
import collections
print('group counts', collections.Counter(zip(d['y'].tolist(), d['g'].tolist())))
r=np.load('data/reweight.npz'); print('reweight', {k:(r[k].shape,r[k].dtype) for k in r})
print('rw group counts', collections.Counter(zip(r['y'].tolist(), r['g'].tolist())))
"train {'X': ((2400, 3, 16, 16), dtype('float32')), 'y': ((2400,), dtype('int64')), 'g': ((2400,), dtype('int64'))}
y vals [0 1 2 3] g vals [0 1 2 3]
group counts Counter({(2, 2): 586, (0, 0): 585, (1, 1): 580, (3, 3): 576, (1, 3): 11, (3, 0): 10, (3, 2): 7, (0, 2): 7, (3, 1): 7, (0, 3): 7, (2, 3): 6, (1, 0): 6, (2, 1): 5, (2, 0): 3, (1, 2): 3, (0, 1): 1})
reweight {'X': ((1280, 3, 16, 16), dtype('float32')), 'y': ((1280,), dtype('int64')), 'g': ((1280,), dtype('int64'))}
rw group counts Counter({(2, 0): 80, (2, 1): 80, (3, 0): 80, (0, 1): 80, (3, 3): 80, (2, 3): 80, (3, 1): 80, (1, 0): 80, (0, 2): 80, (1, 3): 80, (1, 1): 80, (1, 2): 80, (2, 2): 80, (0, 0): 80, (3, 2): 80, (0, 3): 80})/app/submission/autograd.py
TODO: implement the topological walk + seeding. (The per-op ``_backward`` closures do
the local accumulation; here you only orchestrate the order + seed ``self.grad``.)
"""
raise NotImplementedError("Tensor.backward") TODO: implement the topological walk + seeding. (The per-op ``_backward`` closures do
the local accumulation; here you only orchestrate the order + seed ``self.grad``.)
"""
topo: List[Tensor] = []
visited = set()
def build(node: "Tensor"):
if id(node) in visited:
return
visited.add(id(node))
for p in node._parents:
build(p)
topo.append(node)
build(self)
if grad is None:
grad = np.ones_like(self.data)
else:
grad = _as_f64(grad)
self._accum(grad)
for node in reversed(topo):
node._backward()The file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("add backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
if a.requires_grad:
a._accum(_unbroadcast(out.grad, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(out.grad, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("sub backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
if a.requires_grad:
a._accum(_unbroadcast(out.grad, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(-out.grad, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("mul backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
if a.requires_grad:
a._accum(_unbroadcast(out.grad * b.data, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(out.grad * a.data, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("div backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
if a.requires_grad:
a._accum(_unbroadcast(out.grad / b.data, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(-out.grad * a.data / (b.data * b.data), b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("power backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * pf * (a.data ** (pf - 1.0)))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("maximum backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
ad, bd = a.data, b.data
eq = (ad == bd)
ga = np.where(ad > bd, 1.0, 0.0) + 0.5 * eq
gb = np.where(bd > ad, 1.0, 0.0) + 0.5 * eq
if a.requires_grad:
a._accum(_unbroadcast(out.grad * ga, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(out.grad * gb, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("minimum backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
ad, bd = a.data, b.data
eq = (ad == bd)
ga = np.where(ad < bd, 1.0, 0.0) + 0.5 * eq
gb = np.where(bd < ad, 1.0, 0.0) + 0.5 * eq
if a.requires_grad:
a._accum(_unbroadcast(out.grad * ga, a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(out.grad * gb, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("where backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
if a.requires_grad:
a._accum(_unbroadcast(np.where(c, out.grad, 0.0), a.data.shape))
if b.requires_grad:
b._accum(_unbroadcast(np.where(c, 0.0, out.grad), b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("exp backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * e)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("log backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad / a.data)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("sqrt backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * 0.5 / s)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("abs backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * np.sign(a.data))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("clip backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
mask = (a.data > lo) & (a.data < hi)
a._accum(out.grad * mask)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("sum backward") # TODO
out._backward = _bw
return out
def mean(a, axis=None, keepdims=False) -> Tensor: def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
a._accum(np.broadcast_to(g, a.data.shape).copy())
out._backward = _bw
return out
def mean(a, axis=None, keepdims=False) -> Tensor:The file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("mean backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
a._accum(np.broadcast_to(g, a.data.shape) / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("max backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
mask = (a.data == m)
cnt = mask.sum(axis=axis, keepdims=True)
a._accum(mask * g / cnt)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("min backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
mask = (a.data == m)
cnt = mask.sum(axis=axis, keepdims=True)
a._accum(mask * g / cnt)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("transpose backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
if axes is None:
a._accum(np.transpose(out.grad))
else:
inv = np.argsort(axes)
a._accum(np.transpose(out.grad, tuple(inv)))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("reshape backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad.reshape(a.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("getitem backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
gz = np.zeros_like(a.data)
np.add.at(gz, idx, out.grad)
a._accum(gz)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("concat backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
offsets = np.cumsum([0] + sizes)
for i, t in enumerate(ts):
if t.requires_grad:
sl = [slice(None)] * out.grad.ndim
sl[axis] = slice(offsets[i], offsets[i + 1])
t._accum(out.grad[tuple(sl)])
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("stack backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
for i, t in enumerate(ts):
if t.requires_grad:
sl = [slice(None)] * out.grad.ndim
sl[axis] = i
t._accum(out.grad[tuple(sl)])
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("matmul backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
if a.requires_grad:
da = go @ np.swapaxes(b.data, -1, -2)
a._accum(_unbroadcast(da, a.data.shape))
if b.requires_grad:
db = np.swapaxes(a.data, -1, -2) @ go
b._accum(_unbroadcast(db, b.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("relu backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * (a.data > 0.0))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("leaky_relu backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * np.where(a.data > 0.0, 1.0, sl))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("sigmoid backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * s * (1.0 - s))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("tanh backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * (1.0 - t * t))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("gelu backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
pdf = np.exp(-0.5 * x * x) / np.sqrt(2.0 * np.pi)
a._accum(out.grad * (cdf + x * pdf))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("softmax backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
a._accum(s * (g - (g * s).sum(axis=axis, keepdims=True)))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("log_softmax backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
a._accum(g - s * g.sum(axis=axis, keepdims=True))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("softplus backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * (1.0 / (1.0 + np.exp(-bx))))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("silu backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * (sig + a.data * sig * (1.0 - sig)))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("elu backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * np.where(x > 0.0, 1.0, al * ex))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("mish backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
sig = 1.0 / (1.0 + np.exp(-x))
a._accum(out.grad * (t + x * (1.0 - t * t) * sig))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("hardtanh backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
mask = (a.data > lo) & (a.data < hi)
a._accum(out.grad * mask)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("var backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
denom = n - ddof
a._accum(g * (2.0 / denom) * (a.data - mu))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("std backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
if axis is not None and not keepdims:
axes = (axis,) if isinstance(axis, int) else tuple(axis)
axes = tuple(ax % a.data.ndim for ax in axes)
g = np.expand_dims(g, axes)
a._accum(g * (a.data - mu) / (n * skeep))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("cumsum backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
g = out.grad
rev = np.flip(np.cumsum(np.flip(g, axis=axis), axis=axis), axis=axis)
a._accum(rev)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("mse_loss backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not pred.requires_grad:
return
pred._accum(out.grad * 2.0 * diff / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("cross_entropy backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
logits._accum(out.grad * (sm - y) / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("group_weighted_ce backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
g = (sm - y) * scale[:, None]
logits._accum(out.grad * g)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("reweighted_ce backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
g = (sm - y) * (w / wsum)[:, None]
logits._accum(out.grad * g)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("group_dro_loss backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
g = (sm - y) * scale[:, None]
logits._accum(out.grad * g)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("logit_adjusted_ce backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
logits._accum(out.grad * (sm - y) / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("focal_loss backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
omp = 1.0 - p
dfldp = g * (omp ** (g - 1.0)) * np.log(pc) - (omp ** g) / pc
coef = dfldp * p # (N,)
grad = coef[:, None] * (y - sm) / n
logits._accum(out.grad * grad)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("irm_penalty backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
sx = (sm * x).sum(axis=-1, keepdims=True) # <s_i, x_i> (N,1)
dgw = ((sm - y) + sm * (x - sx)) / n
grad = 2.0 * grad_w * dgw
logits._accum(out.grad * grad)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("gce_loss backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
coef = -(pc ** qf) # (N,)
grad = coef[:, None] * (y - sm) / n
logits._accum(out.grad * grad)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("vrex_penalty backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not r.requires_grad:
return
grad = (2.0 / K) * (rd - mu)
r._accum(out.grad * grad.reshape(r.data.shape))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("ldam_loss backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
y = np.zeros_like(sm)
y[np.arange(n), t] = 1.0
logits._accum(out.grad * sc * (sm - y) / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("spectral_decoupling backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not logits.requires_grad:
return
logits._accum(out.grad * lm * x / n)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("layernorm backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
red = tuple(range(go.ndim - 1))
if gamma.requires_grad:
gamma._accum((go * xhat).sum(axis=red))
if beta.requires_grad:
beta._accum(go.sum(axis=red))
if a.requires_grad:
gxhat = go * gamma.data
m1 = gxhat.mean(axis=-1, keepdims=True)
m2 = (gxhat * xhat).mean(axis=-1, keepdims=True)
dx = inv * (gxhat - m1 - xhat * m2)
a._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("batchnorm backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
if gamma.requires_grad:
gamma._accum((go * xhat).sum(axis=0))
if beta.requires_grad:
beta._accum(go.sum(axis=0))
if a.requires_grad:
gxhat = go * gamma.data
m1 = gxhat.mean(axis=0, keepdims=True)
m2 = (gxhat * xhat).mean(axis=0, keepdims=True)
dx = inv * (gxhat - m1 - xhat * m2)
a._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("batchnorm2d backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
if gamma.requires_grad:
gamma._accum((go * xhat).sum(axis=(0, 2, 3)))
if beta.requires_grad:
beta._accum(go.sum(axis=(0, 2, 3)))
if x.requires_grad:
gxhat = go * gamma.data.reshape(1, C, 1, 1)
if training:
m1 = gxhat.sum(axis=(0, 2, 3), keepdims=True) / m
m2 = (gxhat * xhat).sum(axis=(0, 2, 3), keepdims=True) / m
dx = inv * (gxhat - m1 - xhat * m2)
else:
dx = gxhat * inv
x._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("rms_norm backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
if gamma.requires_grad:
red = tuple(range(go.ndim - 1))
gamma._accum((go * xhat).sum(axis=red))
if a.requires_grad:
gy = go * gamma.data
s = (gy * x).sum(axis=-1, keepdims=True)
dx = r * gy - (r ** 3 / D) * x * s
a._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("groupnorm2d backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
go = out.grad
if gamma.requires_grad:
gamma._accum((go * xhat).sum(axis=(0, 2, 3)))
if beta.requires_grad:
beta._accum(go.sum(axis=(0, 2, 3)))
if x.requires_grad:
gxhat = (go * g_).reshape(N, G, m)
xhat_g = xhat.reshape(N, G, m)
m1 = gxhat.sum(axis=2, keepdims=True) / m
m2 = (gxhat * xhat_g).sum(axis=2, keepdims=True) / m
dxg = inv * (gxhat - m1 - xhat_g * m2)
x._accum(dxg.reshape(N, C, H, W))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("conv2d backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None:
return
dop = out.grad.reshape(N, Cout, OH * OW)
if bias.requires_grad:
bias._accum(dop.sum(axis=(0, 2)))
if weight.requires_grad:
dWm = np.einsum("nop,nkp->ok", dop, cols)
weight._accum(dWm.reshape(Cout, Cin, kh, kw))
if x.requires_grad:
dcols = np.einsum("ok,nop->nkp", Wm, dop)
dcols = dcols.reshape(N, Cin, kh, kw, OH, OW)
H, W = x.data.shape[2], x.data.shape[3]
xpg = np.zeros((N, Cin, H + 2 * pad, W + 2 * pad), np.float64)
for i in range(kh):
for j in range(kw):
xpg[:, :, i:i + st * OH:st, j:j + st * OW:st] += dcols[:, :, i, j, :, :]
if pad:
dx = xpg[:, :, pad:pad + H, pad:pad + W]
else:
dx = xpg
x._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("maxpool2d backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not x.requires_grad:
return
mexp = outd[:, :, :, None, :, None]
mask = (xr == mexp)
cnt = mask.sum(axis=(3, 5), keepdims=True)
go = out.grad[:, :, :, None, :, None]
gxr = mask * go / cnt
x._accum(gxr.reshape(N, C, H, W))
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("maxpool2d_stride backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not x.requires_grad:
return
mexp = outd[:, :, None, None, :, :]
mask = (win == mexp)
cnt = mask.sum(axis=(2, 3), keepdims=True)
go = out.grad[:, :, None, None, :, :]
gwin = mask * go / cnt
xpg = np.zeros((N, C, Hp, Wp), np.float64)
for i in range(k):
for j in range(k):
xpg[:, :, i:i + st * OH:st, j:j + st * OW:st] += gwin[:, :, i, j, :, :]
if pad:
dx = xpg[:, :, pad:pad + H, pad:pad + W]
else:
dx = xpg
x._accum(dx)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("avgpool2d backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not x.requires_grad:
return
go = out.grad[:, :, :, None, :, None]
gxr = np.broadcast_to(go / (k * k), (N, C, H // k, k, W // k, k))
x._accum(gxr.reshape(N, C, H, W).copy())
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("global_avg_pool backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not x.requires_grad:
return
go = out.grad[:, :, None, None]
x._accum(np.broadcast_to(go / (H * W), (N, C, H, W)).copy())
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("embedding backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not weight.requires_grad:
return
gz = np.zeros_like(weight.data)
np.add.at(gz, idx, out.grad)
weight._accum(gz)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw_id():
raise NotImplementedError("dropout backward") # TODO
out._backward = _bw_id
return out def _bw_id():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad)
out._backward = _bw_id
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/autograd.py
def _bw():
raise NotImplementedError("dropout backward") # TODO
out._backward = _bw
return out def _bw():
if out.grad is None or not a.requires_grad:
return
a._accum(out.grad * keep)
out._backward = _bw
return outThe file /app/submission/autograd.py has been updated successfully. (file state is current in your context , no need to Read it back)
Run autograd grad checks
OMP_NUM_THREADS=1 python3 run_checks.py --only autograd_grads 2>&1 | tail -40
grad/hardtanh OK (max_abs=1.57e-11) grad/var_axis OK (max_abs=3.14e-11) grad/std_axis OK (max_abs=3.54e-11) grad/cumsum OK (max_abs=1.14e-10) grad/maxpool2d_stride OK (max_abs=7.86e-10) grad/layernorm_x OK grad/batchnorm1d_x OK grad/embedding OK grad/mse_loss OK grad/cross_entropy OK grad/conv2d_x OK grad/conv2d_weight OK grad/conv2d_bias OK grad/conv2d_x_pad0 OK grad/batchnorm2d_x OK grad/batchnorm2d_gamma OK grad/batchnorm2d_beta OK grad/batchnorm2d_x_eval OK grad/nn.Linear_x FAIL (raised NotImplementedError: forward) grad/nn.Conv2d_x FAIL (raised NotImplementedError: forward) grad/nn.LayerNorm_x FAIL (raised NotImplementedError: forward) grad/nn.BatchNorm1d_x FAIL (raised NotImplementedError: forward) grad/CNN_grad_classifier.weight FAIL (raised NotImplementedError: forward) grad/CNN_grad_embed.weight FAIL (rai …[truncated 1123 chars]
Check for non-forward failures
OMP_NUM_THREADS=1 python3 run_checks.py --only autograd_grads 2>&1 | grep -c "OK"; OMP_NUM_THREADS=1 python3 run_checks.py --only autograd_grads 2>&1 | grep "FAIL" | grep -v "forward"
67
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
d = {"weight": self.weight} def forward(self, x: Tensor) -> Tensor:
out = ag.matmul(x, self.weight)
if self.bias is not None:
out = ag.add(out, self.bias)
return out
def named_params(self):
d = {"weight": self.weight}The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"weight": self.weight, "bias": self.bias} def forward(self, x: Tensor) -> Tensor:
return ag.conv2d(x, self.weight, self.bias, pad=self.pad, stride=1)
def named_params(self):
return {"weight": self.weight, "bias": self.bias}The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor, training: bool = True) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
def buffers(self): def forward(self, x: Tensor, training: bool = True) -> Tensor:
return ag.batchnorm2d(x, self.gamma, self.beta,
self.running_mean, self.running_var,
training=training, momentum=self.momentum, eps=self.eps)
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
def buffers(self):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
class LayerNorm(Module): def forward(self, x: Tensor) -> Tensor:
return ag.batchnorm(x, self.gamma, self.beta, eps=self.eps)
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
class LayerNorm(Module):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
class Dropout(Module): def forward(self, x: Tensor) -> Tensor:
return ag.layernorm(x, self.gamma, self.beta, eps=self.eps)
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}
class Dropout(Module):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor, training: bool = False, rng=None) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {}
class AvgPool2d(Module): def forward(self, x: Tensor, training: bool = False, rng=None) -> Tensor:
return ag.dropout(x, self.p, training, rng)
def named_params(self):
return {}
class AvgPool2d(Module):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {}
class MaxPool2d(Module): def forward(self, x: Tensor) -> Tensor:
return ag.avgpool2d(x, self.k)
def named_params(self):
return {}
class MaxPool2d(Module):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {}
class RMSNorm(Module): def forward(self, x: Tensor) -> Tensor:
return ag.maxpool2d(x, self.k)
def named_params(self):
return {}
class RMSNorm(Module):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.rms_norm). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"gamma": self.gamma} def forward(self, x: Tensor) -> Tensor:
return ag.rms_norm(x, self.gamma, eps=self.eps)
def named_params(self):
return {"gamma": self.gamma}The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x: Tensor) -> Tensor:
# TODO: implement (compose ag.groupnorm2d). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta} def forward(self, x: Tensor) -> Tensor:
return ag.groupnorm2d(x, self.gamma, self.beta, self.num_groups, eps=self.eps)
def named_params(self):
return {"gamma": self.gamma, "beta": self.beta}The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, X, training: bool = True) -> Tensor:
# TODO: implement (stem->gn0->silu; then 2 residual blocks
# r=h; h=silu(b*n1(b*c1(h))); h=b*n2(b*c2(h)); h=silu(h + r);
# then strided downsample via ag.conv2d(h, down.weight, down.bias, pad=1, stride=2) + silu;
# then global-avg-pool over axes (2,3); then the linear head). Compose ag.* ops.
raise NotImplementedError("forward") def forward(self, X, training: bool = True) -> Tensor:
h = X if isinstance(X, Tensor) else Tensor(X)
h = ag.silu(self.gn0(self.stem(h)))
# block 1
r = h
h = ag.silu(self.b1n1(self.b1c1(h)))
h = self.b1n2(self.b1c2(h))
h = ag.silu(ag.add(h, r))
# block 2
r = h
h = ag.silu(self.b2n1(self.b2c1(h)))
h = self.b2n2(self.b2c2(h))
h = ag.silu(ag.add(h, r))
# strided downsample
h = ag.silu(ag.conv2d(h, self.down.weight, self.down.bias, pad=1, stride=2))
h = ag.global_avg_pool(h)
return self.head(h)The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def featurize(self, X, training: bool = True) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("featurize")
def forward(self, X, training: bool = True) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward") def featurize(self, X, training: bool = True) -> Tensor:
h = X if isinstance(X, Tensor) else Tensor(X)
h = ag.maxpool2d(ag.relu(self.bn1(self.conv1(h), training=training)), 2)
h = ag.maxpool2d(ag.relu(self.bn2(self.conv2(h), training=training)), 2)
h = ag.relu(self.bn3(self.conv3(h), training=training))
h = ag.global_avg_pool(h)
h = ag.relu(self.embed(h))
return h
def forward(self, X, training: bool = True) -> Tensor:
return self.classifier(self.featurize(X, training=training))The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for i, lin in enumerate(self.layers): def forward(self, x) -> Tensor:
h = x if isinstance(x, Tensor) else Tensor(x)
for i, lin in enumerate(self.layers):
h = lin(h)
if i < len(self.layers) - 1:
h = ag.relu(h)
return h
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for i, lin in enumerate(self.layers):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, x) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for k, v in self.proj_in.named_params().items():
d[f"proj_in.{k}"] = v def forward(self, x) -> Tensor:
h = x if isinstance(x, Tensor) else Tensor(x)
h = self.proj_in(h)
for blk in self.blocks:
y = blk["ln"](h)
y = blk["fc1"](y)
y = ag.gelu(y)
y = blk["fc2"](y)
h = ag.add(h, y)
return self.proj_out(h)
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for k, v in self.proj_in.named_params().items():
d[f"proj_in.{k}"] = vThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def forward(self, X, training: bool = True) -> Tensor:
# TODO: implement (compose ag.* ops). See the docstring/contract.
raise NotImplementedError("forward")
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for name in ("conv1", "bn1", "conv2", "bn2", "head"): def forward(self, X, training: bool = True) -> Tensor:
h = X if isinstance(X, Tensor) else Tensor(X)
h = ag.avgpool2d(ag.relu(self.bn1(self.conv1(h), training=training)), 2)
h = ag.maxpool2d(ag.relu(self.bn2(self.conv2(h), training=training)), 2)
N = h.data.shape[0]
h = ag.reshape(h, (N, self.W * 2 * self.side * self.side))
return self.head(h)
def named_params(self) -> Dict[str, Tensor]:
d: Dict[str, Tensor] = {}
for name in ("conv1", "bn1", "conv2", "bn2", "head"):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
balanced-group / DFR weighting). If ``n_groups`` is None it is inferred as max(group_ids)+1."""
# TODO: implement (numpy only). See docstring.
raise NotImplementedError("balanced_group_weights") balanced-group / DFR weighting). If ``n_groups`` is None it is inferred as max(group_ids)+1."""
gids = np.asarray(group_ids).astype(np.int64).reshape(-1)
if n_groups is None:
n_groups = int(gids.max()) + 1
present = np.unique(gids)
w = np.zeros(int(n_groups), dtype=np.float64)
w[present] = 1.0 / len(present)
return wThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
effective number is (1-beta^n_c)/(1-beta) (Cui et al.); beta=0 reduces to inverse-frequency
1/n_c. Each sample gets its class weight, and the vector is normalized to mean 1."""
# TODO: implement (numpy only). See docstring.
raise NotImplementedError("class_balanced_weights") effective number is (1-beta^n_c)/(1-beta) (Cui et al.); beta=0 reduces to inverse-frequency
1/n_c. Each sample gets its class weight, and the vector is normalized to mean 1."""
t = np.asarray(targets).astype(np.int64).reshape(-1)
if n_classes is None:
n_classes = int(t.max()) + 1
counts = np.bincount(t, minlength=int(n_classes)).astype(np.float64)
cw = np.zeros(int(n_classes), dtype=np.float64)
nz = counts > 0
if beta == 0.0:
cw[nz] = 1.0 / counts[nz]
else:
eff = (1.0 - beta ** counts[nz]) / (1.0 - beta)
cw[nz] = 1.0 / eff
w = cw[t]
w = w / w.mean()
return wThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
vector. Empty classes get log(eps)."""
# TODO: implement (numpy only). See docstring.
raise NotImplementedError("log_class_prior") vector. Empty classes get log(eps)."""
t = np.asarray(targets).astype(np.int64).reshape(-1)
if n_classes is None:
n_classes = int(t.max()) + 1
counts = np.bincount(t, minlength=int(n_classes)).astype(np.float64)
N = t.size
prior = counts / N
out = np.where(prior > 0, np.log(np.maximum(prior, eps)), np.log(eps))
return outThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
Returns a length-``n_classes`` vector."""
# TODO: implement (numpy only). See docstring.
raise NotImplementedError("ldam_margins") Returns a length-``n_classes`` vector."""
t = np.asarray(targets).astype(np.int64).reshape(-1)
if n_classes is None:
n_classes = int(t.max()) + 1
counts = np.bincount(t, minlength=int(n_classes)).astype(np.float64)
m = np.zeros(int(n_classes), dtype=np.float64)
nz = counts > 0
m[nz] = counts[nz] ** (-0.25)
mx = m[nz].max() if nz.any() else 1.0
m[nz] = m[nz] * (max_margin / mx)
return mThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
the PRESENT classes. Returns a length-``n_classes`` per-CLASS weight vector (NOT per-sample)."""
# TODO: implement (numpy only). See docstring.
raise NotImplementedError("effective_number_weights") the PRESENT classes. Returns a length-``n_classes`` per-CLASS weight vector (NOT per-sample)."""
t = np.asarray(targets).astype(np.int64).reshape(-1)
if n_classes is None:
n_classes = int(t.max()) + 1
counts = np.bincount(t, minlength=int(n_classes)).astype(np.float64)
w = np.zeros(int(n_classes), dtype=np.float64)
nz = counts > 0
eff = (1.0 - beta ** counts[nz]) / (1.0 - beta)
w[nz] = 1.0 / eff
w[nz] = w[nz] / w[nz].mean()
return wThe file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def update(self, group_losses, present=None):
# TODO: implement the masked EMA update (see docstring). Return self.l.copy().
raise NotImplementedError("EMAGroupLoss.update")
def worst_group(self):
# TODO: implement (argmax of the current EMA). See docstring.
raise NotImplementedError("EMAGroupLoss.worst_group") def update(self, group_losses, present=None):
gl = np.asarray(group_losses, dtype=np.float64).reshape(-1)
if present is None:
mask = np.ones(self.n_groups, dtype=bool)
else:
mask = np.asarray(present, dtype=bool).reshape(-1)
for k in range(self.n_groups):
if not mask[k]:
continue
if not self._init[k]:
self.l[k] = gl[k]
self._init[k] = True
else:
self.l[k] = self.beta * self.l[k] + (1.0 - self.beta) * gl[k]
return self.l.copy()
def worst_group(self):
return int(np.argmax(self.l))The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/nn.py
def update(self, group_losses):
# TODO: exponentiated-gradient update of self.q on group_losses; renormalize.
raise NotImplementedError("GroupDROState.update")
def weights(self): def update(self, group_losses):
gl = np.asarray(group_losses, dtype=np.float64).reshape(-1)
self.q = self.q * np.exp(self.eta_q * gl)
self.q = self.q / self.q.sum()
return self.q.copy()
def weights(self):The file /app/submission/nn.py has been updated successfully. (file state is current in your context , no need to Read it back)
Run nn/robust checks
OMP_NUM_THREADS=1 python3 run_checks.py --only nn_forward 2>&1 | tail -8; echo "===ROBUST==="; OMP_NUM_THREADS=1 python3 run_checks.py --only robust 2>&1 | tail -5; echo "===ROBUST2==="; OMP_NUM_THREADS=1 python3 run_checks.py --only robust2 2>&1 | tail -5
fwd/MLP_shape OK fwd/ResidualMLP_shape OK fwd/TinyConvNet_shape OK ---------------------------------------------------------------------- nn_forward 21/21 TOTAL 21/21 public checks passing All public checks pass. Now make sure model.npz clears the held-out worst-group accuracy threshold; the grader re-checks with DIFFERENT seeds. ===ROBUST=== ---------------------------------------------------------------------- robust 13/13 TOTAL 13/13 public checks passing All public checks pass. Now make sure model.npz clears the held-out worst-group accuracy threshold; the grader re-checks with DIFFERENT seeds. ===ROBUST2=== ---------------------------------------------------------------------- robust2 8/8 TOTAL 8/8 public checks passing All public checks pass. Now make sure model.npz clears the held-out worst-group accuracy threshold; the grader re-checks with DIFFERENT seeds.
/app/submission/optim.py
def step(self):
# TODO: implement the SGD update (see module docstring). Update p.data in place.
raise NotImplementedError("SGD.step") def step(self):
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._v[i] = self.momentum * self._v[i] + g
v = self._v[i]
upd = (g + self.momentum * v) if self.nesterov else v
p.data -= self.lr * updThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the Adagrad update (see module docstring). Update p.data in place.
raise NotImplementedError("Adagrad.step") def step(self):
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._s[i] = self._s[i] + g * g
p.data -= self.lr * g / (np.sqrt(self._s[i]) + self.eps)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the RMSprop update (see module docstring). Update p.data in place.
raise NotImplementedError("RMSprop.step") def step(self):
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._s[i] = self.alpha * self._s[i] + (1 - self.alpha) * (g * g)
p.data -= self.lr * g / (np.sqrt(self._s[i]) + self.eps)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the Adam update (see module docstring). Update p.data in place.
raise NotImplementedError("Adam.step") def step(self):
self.t += 1
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._m[i] = self.b1 * self._m[i] + (1 - self.b1) * g
self._v[i] = self.b2 * self._v[i] + (1 - self.b2) * (g * g)
mhat = self._m[i] / (1 - self.b1 ** self.t)
vhat = self._v[i] / (1 - self.b2 ** self.t)
p.data -= self.lr * mhat / (np.sqrt(vhat) + self.eps)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the AdamW update (see module docstring). Update p.data in place.
raise NotImplementedError("AdamW.step") def step(self):
self.t += 1
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad
self._m[i] = self.b1 * self._m[i] + (1 - self.b1) * g
self._v[i] = self.b2 * self._v[i] + (1 - self.b2) * (g * g)
mhat = self._m[i] / (1 - self.b1 ** self.t)
vhat = self._v[i] / (1 - self.b2 ** self.t)
p.data -= self.lr * (mhat / (np.sqrt(vhat) + self.eps) + self.weight_decay * p.data)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the Adamax update (see module docstring). Update p.data in place.
raise NotImplementedError("Adamax.step") def step(self):
self.t += 1
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._m[i] = self.b1 * self._m[i] + (1 - self.b1) * g
self._u[i] = np.maximum(self.b2 * self._u[i], np.abs(g))
p.data -= (self.lr / (1 - self.b1 ** self.t)) * self._m[i] / (self._u[i] + self.eps)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the NAdam update (see module docstring). Update p.data in place.
raise NotImplementedError("NAdam.step") def step(self):
self.t += 1
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._m[i] = self.b1 * self._m[i] + (1 - self.b1) * g
self._v[i] = self.b2 * self._v[i] + (1 - self.b2) * (g * g)
mhat = self._m[i] / (1 - self.b1 ** self.t)
vhat = self._v[i] / (1 - self.b2 ** self.t)
mnes = self.b1 * mhat + (1 - self.b1) * g / (1 - self.b1 ** self.t)
p.data -= self.lr * mnes / (np.sqrt(vhat) + self.eps)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the Lion update (see module docstring). Update p.data in place.
raise NotImplementedError("Lion.step") def step(self):
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad
upd = np.sign(self.b1 * self._m[i] + (1 - self.b1) * g)
p.data -= self.lr * (upd + self.weight_decay * p.data)
self._m[i] = self.b2 * self._m[i] + (1 - self.b2) * gThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the RAdam update (see module docstring). Update p.data in place.
raise NotImplementedError("RAdam.step") def step(self):
self.t += 1
t = self.t
b1, b2 = self.b1, self.b2
bc2 = 1 - b2 ** t
rinf = 2.0 / (1 - b2) - 1
rt = rinf - 2 * t * (b2 ** t) / bc2
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad + self.weight_decay * p.data
self._m[i] = b1 * self._m[i] + (1 - b1) * g
self._v[i] = b2 * self._v[i] + (1 - b2) * (g * g)
mhat = self._m[i] / (1 - b1 ** t)
if rt > 4:
l = np.sqrt(bc2 / (self._v[i] + self.eps))
r = math.sqrt(((rt - 4) * (rt - 2) * rinf) / ((rinf - 4) * (rinf - 2) * rt))
p.data -= self.lr * mhat * r * l
else:
p.data -= self.lr * mhatThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
def step(self):
# TODO: implement the AdaBelief update (see module docstring). Update p.data in place.
raise NotImplementedError("AdaBelief.step") def step(self):
self.t += 1
for i, p in enumerate(self.params):
if p.grad is None:
continue
g = p.grad
self._m[i] = self.b1 * self._m[i] + (1 - self.b1) * g
d = g - self._m[i]
self._s[i] = self.b2 * self._s[i] + (1 - self.b2) * (d * d) + self.eps
mhat = self._m[i] / (1 - self.b1 ** self.t)
shat = self._s[i] / (1 - self.b2 ** self.t)
p.data -= self.lr * (mhat / (np.sqrt(shat) + self.eps) + self.weight_decay * p.data)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
exceeds ``max_norm`` every grad is scaled by ``max_norm / (total_norm + 1e-6)``."""
# TODO: implement (see docstring).
raise NotImplementedError("clip_grad_norm") exceeds ``max_norm`` every grad is scaled by ``max_norm / (total_norm + 1e-6)``."""
total = 0.0
for p in params:
if p.grad is not None:
total += float((p.grad ** 2).sum())
total = math.sqrt(total)
if total > max_norm:
scale = max_norm / (total + 1e-6)
for p in params:
if p.grad is not None:
p.grad = p.grad * scale
return totalThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
"""Element-wise gradient clamp to [-clip_value, clip_value] (in place)."""
# TODO: implement (see docstring).
raise NotImplementedError("clip_grad_value") """Element-wise gradient clamp to [-clip_value, clip_value] (in place)."""
for p in params:
if p.grad is not None:
p.grad = np.clip(p.grad, -clip_value, clip_value)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
``total``. ``step`` is 0-based. Holds at ``min_lr`` past ``total``."""
# TODO: implement (see docstring).
raise NotImplementedError("cosine_lr") ``total``. ``step`` is 0-based. Holds at ``min_lr`` past ``total``."""
if warmup > 0 and step < warmup:
return base_lr * (step + 1) / warmup
prog = min(1.0, max(0.0, (step - warmup) / max(1, total - warmup)))
return min_lr + 0.5 * (base_lr - min_lr) * (1 + math.cos(math.pi * prog))The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
"""Linear warmup to ``base_lr`` over ``warmup`` steps, then constant ``base_lr``."""
# TODO: implement (see docstring).
raise NotImplementedError("warmup_lr") """Linear warmup to ``base_lr`` over ``warmup`` steps, then constant ``base_lr``."""
if step < warmup:
return base_lr * (step + 1) / warmup
return base_lrThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
"""Step decay: lr = base_lr * gamma ** (step // step_size)."""
# TODO: implement (see docstring).
raise NotImplementedError("step_lr") """Step decay: lr = base_lr * gamma ** (step // step_size)."""
return base_lr * (gamma ** (step // step_size))The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
"""Exponential decay: lr = base_lr * gamma ** step."""
# TODO: implement (see docstring).
raise NotImplementedError("exponential_lr") """Exponential decay: lr = base_lr * gamma ** step."""
return base_lr * (gamma ** step)The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
``step >= total`` holds at ``min_lr``. ``step`` is 0-based."""
# TODO: implement (see docstring).
raise NotImplementedError("polynomial_lr") ``step >= total`` holds at ``min_lr``. ``step`` is 0-based."""
if step >= total:
return min_lr
return (base_lr - min_lr) * ((1 - step / total) ** power) + min_lrThe file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/optim.py
in-cycle step s: lr = min_lr + 0.5*(base_lr-min_lr)*(1+cos(pi*s/T)). ``step`` is 0-based."""
# TODO: implement (see docstring).
raise NotImplementedError("cosine_restart_lr") in-cycle step s: lr = min_lr + 0.5*(base_lr-min_lr)*(1+cos(pi*s/T)). ``step`` is 0-based."""
s = step
T_ = period
while s >= T_:
s -= T_
T_ *= t_mult
return min_lr + 0.5 * (base_lr - min_lr) * (1 + math.cos(math.pi * s / max(1, T_)))The file /app/submission/optim.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/data.py
"""Flatten (y, g) to a single group id y*n_classes + g (int64)."""
# TODO
raise NotImplementedError("group_ids") """Flatten (y, g) to a single group id y*n_classes + g (int64)."""
return (np.asarray(y).astype(np.int64) * int(n_classes) + np.asarray(g).astype(np.int64))The file /app/submission/data.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/data.py
"""Yield (Xb, yb, gb) minibatches over a random permutation of the dataset."""
# TODO
raise NotImplementedError("iterate") """Yield (Xb, yb, gb) minibatches over a random permutation of the dataset."""
N = X.shape[0]
perm = rng.permutation(N)
for i in range(0, N, batch_size):
idx = perm[i:i + batch_size]
yield X[idx], y[idx], g[idx]The file /app/submission/data.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/data.py
"""Return (min over (y,g) groups of within-group accuracy, average accuracy, per-group dict)."""
# TODO
raise NotImplementedError("worst_group_acc") """Return (min over (y,g) groups of within-group accuracy, average accuracy, per-group dict)."""
preds = np.asarray(preds).reshape(-1)
y = np.asarray(y).reshape(-1)
g = np.asarray(g).reshape(-1)
gid = group_ids(y, g, n_classes)
correct = (preds == y)
avg = float(correct.mean())
per = {}
for k in np.unique(gid):
mask = gid == k
per[int(k)] = float(correct[mask].mean())
wga = min(per.values()) if per else 0.0
return wga, avg, perThe file /app/submission/data.py has been updated successfully. (file state is current in your context , no need to Read it back)
/app/submission/data.py
"""Boolean mask selecting ~frac of EACH (y,g) group (keeps both slices group-balanced)."""
# TODO
raise NotImplementedError("stratified_group_split") """Boolean mask selecting ~frac of EACH (y,g) group (keeps both slices group-balanced)."""
y = np.asarray(y).reshape(-1)
g = np.asarray(g).reshape(-1)
gid = group_ids(y, g, n_classes)
mask = np.zeros(len(y), dtype=bool)
for k in np.unique(gid):
idx = np.where(gid == k)[0]
rng.shuffle(idx)
n_sel = int(round(frac * len(idx)))
mask[idx[:n_sel]] = True
return maskThe file /app/submission/data.py has been updated successfully. (file state is current in your context , no need to Read it back)
Run full check suite
OMP_NUM_THREADS=1 python3 run_checks.py 2>&1 | tail -12
robust2/effective_number_weights OK (mean=1.000) robust2/ema_group_loss_tracks_worst OK (l=[0.112 0.904 0.196]) [data] 3/3 data/group_ids OK ([ 1 4 11 14]) data/worst_group_acc OK (wga=0.00 avg=0.80) data/stratified_group_split OK (selected 33/60) ---------------------------------------------------------------------- autograd_grads 81/81 | nn_forward 21/21 | optim 21/21 | robust 13/13 | robust2 8/8 | data 3/3 TOTAL 147/147 public checks passing All public checks pass. Now make sure model.npz clears the held-out worst-group accuracy threshold; the grader re-checks with DIFFERENT seeds.
View test.sh
cat test.sh
#!/usr/bin/env bash
# Smoke test (NOT the grader): confirms the environment is wired -- numpy imports, the train +
# reweight splits load, and the submission stubs are present. It does NOT check correctness (the
# stubs raise NotImplementedError until you implement them).
set -uo pipefail
python3 - <<'PY'
import numpy as np, os
for f in ("train.npz", "reweight.npz"):
assert os.path.exists(f"/app/data/{f}"), f"missing /app/data/{f}"
d = np.load("/app/data/train.npz"); r = np.load("/app/data/reweight.npz")
print(f"train X {d['X'].shape} y {d['y'].shape} g {d['g'].shape}; reweight X {r['X'].shape}")
assert r["X"].shape[1:] == d["X"].shape[1:], (r["X"].shape, d["X"].shape)
for f in ("autograd.py", "nn.py", "optim.py", "data.py", "train.py"):
assert os.path.exists(f"/app/submission/{f}"), f"missing /app/submission/{f}"
assert os.path.exists("/app/run_checks.py"), "missing /app/run_checks.py"
print("submission stubs present:", sorted(os.listdir("/app/submission")))
print("numpy", np.__version__)
PY
# The visible check runner is present + runnable (it exits 1 on the unfilled stub, which is
# expected -- the stubs raise NotImplementedError until you implement them).
python3 /app/run_checks.py --quiet || true
echo "app_test smoke OK (run 'python3 /app/run_checks.py' to test your implementation)"/app/submission/train.py
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--train", default="/app/data/train.npz")
ap.add_argument("--reweight", default="/app/data/reweight.npz")
ap.add_argument("--out", default="/app/submission/model.npz")
ap.add_argument("--seed", type=int, default=0)
args = ap.parse_args()
# TODO: build a SmallCNN(n_classes, img_size, width=WIDTH, embed_dim=EMBED_DIM), train a
# worst-group-robust recipe with your framework (forward -> loss -> backward -> clip ->
# optimizer step), then save:
# save_checkpoint(args.out, model.state_dict(), n_classes, img_size)
raise NotImplementedError("implement training and call save_checkpoint")def _featurize_np(model, X, batch=256):
"""Compute penultimate features (N, E) in eval mode, batched, as a numpy array."""
feats = []
for i in range(0, X.shape[0], batch):
xb = X[i:i + batch]
f = model.featurize(xb, training=False)
feats.append(np.asarray(f.data))
return np.concatenate(feats, axis=0)
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--train", default="/app/data/train.npz")
ap.add_argument("--reweight", default="/app/data/reweight.npz")
ap.add_argument("--out", default="/app/submission/model.npz")
ap.add_argument("--seed", type=int, default=0)
args = ap.parse_args()
rng = np.random.default_rng(args.seed)
Xtr, ytr, gtr = D.load_npz(args.train)
Xrw, yrw, grw = D.load_npz(args.reweight)
n_classes = int(max(ytr.max(), yrw.max())) + 1
img_size = Xtr.shape[-1]
print(f"n_classes={n_classes} img_size={img_size} Xtr={Xtr.shape} Xrw={Xrw.shape}")
model = SmallCNN(n_classes=n_classes, img_size=img_size,
width=WIDTH, embed_dim=EMBED_DIM, seed=args.seed)
# ---------------------------------------------------------------- Phase 1: ERM extractor
params = model.params()
epochs = 25
batch = 128
base_lr = 2e-3
n = Xtr.shape[0]
steps_per_epoch = (n + batch - 1) // batch
total_steps = epochs * steps_per_epoch
opt = O.AdamW(params, lr=base_lr, weight_decay=1e-4)
step = 0
for ep in range(epochs):
perm = r
…[truncated 3593 chars]The file /app/submission/train.py has been updated successfully. (file state is current in your context , no need to Read it back)
Run training
cd /app/submission && OMP_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 MKL_NUM_THREADS=1 timeout 580 python3 train.py 2>&1 | tail -40
No file changes, the agent left the workspace untouched.
=== SyncVals verifier (local re-grade of the agent's own checkpoint) === COMPONENT CONFORMANCE SUITE: 194/194 cases passed (all milestones pass: True) FINAL METRIC: worst-group accuracy = 0.8067 < threshold 0.86 REWARD = 0 (1 iff every component milestone passes AND the sealed held-out metric clears its threshold) NOTE: the in-sandbox training in the 481MB e2b box was OOM-killed (0-byte checkpoint), so the raw in-sandbox reward was an infrastructure artifact. This reward is from a faithful LOCAL re-train of the agent's OWN train.py (24GB, no OOM) with its own default hyperparameters, then re-graded by the real verifier.
Reproduce this trial: git checkout 2f94510 && PYTHONPATH=src python3 scripts/build_site.py , then open trial/trial_d266b06a6b214868. Re-running the agent live requires EVAL_PLATFORM_ENABLE_OAUTH_SMOKE=1 and is non-deterministic.
Trial trial_d266b06a6b214868 · verifier authoritative; classifier explanatory.