Skip to content

Commit 5047409

Browse files
committed
feat(ci): auto-apply triage labels to issues
Issues created via the API or gh bypass the issue templates and arrive without any triage/* label; roughly half of the open issues carry none. Label issues on arrival (opened/reopened) and run a daily sweep as a backstop: issues already prioritised, assigned, epics, or frozen get triage/accepted, the rest triage/needs-triage. The sweep paces its writes and caps how many it makes per run, so a pass over a large backlog stays inside GitHub's secondary rate limits instead of failing partway through. The manual dispatch defaults to dry-run, so the one entry point a human can reach writes nothing until somebody unchecks the box. The schedule is the one event that writes without being asked. No PR lane can exercise this workflow, so the first real run is a sweep over every open issue. A contract test pins the executable lines that bound that: the pull-request skip, the already-triaged skip, the pacing and the cap, the dry-run default, the pinned action, and the token scopes. Assisted-By: Claude <[email protected]> Signed-off-by: Aleksei Sviridkin <[email protected]>
1 parent 677d80a commit 5047409

2 files changed

Lines changed: 618 additions & 0 deletions

File tree

Lines changed: 288 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,288 @@
1+
name: Issue triage
2+
3+
# Every open issue should carry a triage/* label marking its lifecycle stage.
4+
# The issue templates apply `triage/needs-triage` from their front matter, but
5+
# that only covers issues created through the web form: issues created via
6+
# the API or `gh issue create` bypass templates and arrive unlabeled. The
7+
# `issues` events below label those on arrival; the daily sweep is the
8+
# backstop for removed labels and for the pre-existing backlog.
9+
on:
10+
issues:
11+
types: [opened, reopened]
12+
schedule:
13+
- cron: '53 5 * * *' # daily at 05:53 UTC, after stale (04:37)
14+
workflow_dispatch:
15+
inputs:
16+
dry_run:
17+
description: 'Log what would be labeled without applying'
18+
type: boolean
19+
# On by default, so the one entry point a human can reach writes
20+
# nothing until somebody unchecks it. The scheduled sweep reads no
21+
# inputs and is unaffected: it writes.
22+
default: true
23+
24+
# Top-level token defaults to read-only. A job-level block replaces this one
25+
# rather than adding to it, so the job below holds `issues: write` and
26+
# nothing else.
27+
permissions:
28+
contents: read
29+
30+
# One queue for every run of this workflow, on purpose. GitHub keeps a
31+
# single pending run per group, so issues filed while a sweep is running
32+
# lose all but the last of their own runs and wait for the next sweep,
33+
# which is what the sweep is there for. Splitting the group per issue would
34+
# save that wait at the price of a real bug: a sweep works from a listing
35+
# taken when it started, so it and a concurrent per-issue run can read the
36+
# same issue differently and leave it holding two contradictory triage
37+
# labels, which nothing later removes.
38+
#
39+
# The queue closes that door against a second run of this workflow, not
40+
# against a maintainer labeling an issue by hand while a sweep is running:
41+
# the sweep classifies from the listing it took when it started, so a label
42+
# applied during its paced writes is invisible to it and the issue ends up
43+
# carrying a second triage label that the already-triaged skip then keeps
44+
# every later sweep from revisiting. Losing that race costs one label
45+
# somebody removes by hand, which was judged cheaper than re-reading every
46+
# issue immediately before every write: that would double the request
47+
# budget the write cap below is derived against.
48+
concurrency:
49+
group: issue-triage
50+
cancel-in-progress: false
51+
52+
jobs:
53+
triage:
54+
runs-on: ubuntu-latest
55+
# A run that reaches the write cap paces roughly seven minutes of work, so
56+
# anything near this ceiling is a call that hung rather than a long sweep.
57+
# Without it the default is six hours, and the queue above means one hung
58+
# run holds every issue filed behind it for all six.
59+
timeout-minutes: 20
60+
permissions:
61+
issues: write
62+
steps:
63+
- name: Apply triage labels
64+
uses: actions/github-script@f28e40c7f34bde8b3046d885e986cb6290c5673b # v7
65+
with:
66+
script: |
67+
// Labels applied with the workflow's GITHUB_TOKEN do not trigger
68+
// other workflows (GitHub's recursion prevention), so this job
69+
// cannot re-enter itself or wake any `labeled`-subscribed workflow.
70+
//
71+
// Adding a label bumps the issue's `updated_at`, which stale.yaml
72+
// reads as fresh activity: the stale bot un-marks `lifecycle/stale`
73+
// and restarts its 60-day clock on the next run. For issues labeled
74+
// on arrival that is moot; for the sweep over the historical
75+
// backlog it is a deliberate one-time reset. Triage first, then
76+
// let staleness take its course.
77+
//
78+
// `triage/accepted` is also on stale.yaml's exempt-issue-labels,
79+
// so applying it is more than a clock reset: the issue stops
80+
// being auto-closable at all. For every label in
81+
// ACCEPTED_SIGNAL_LABELS that changes nothing, because each one
82+
// is on the same exempt list already. An assignee is the single
83+
// signal that is not a label and not already exempt, so an issue
84+
// whose only accepted-signal is an assignee gains something it
85+
// did not have. That is new, and it is the
86+
// intended reading rather than a side effect: somebody took
87+
// ownership, which makes the issue a known long-tail task and
88+
// not an abandoned one — the same argument stale.yaml's own
89+
// comment makes for exempting `triage/accepted`. Part of the
90+
// backlog this reaches is marked stale today, some of it inside
91+
// the 14-day close window, and the sweep spares it.
92+
const NEEDS_TRIAGE = 'triage/needs-triage';
93+
const ACCEPTED = 'triage/accepted';
94+
95+
// Label-shaped accepted-signals. Every entry is on stale.yaml's
96+
// exempt-issue-labels, and that is the point rather than a
97+
// coincidence: `triage/accepted` is itself exempt, so a signal
98+
// that is not already exempt would hand the issue a permanent
99+
// reprieve as a side effect of being labeled. `priority/backlog`
100+
// is the tier stale.yaml deliberately left out, so it is left out
101+
// here too and those issues go through triage like any other.
102+
const ACCEPTED_SIGNAL_LABELS = [
103+
'priority/critical-urgent',
104+
'priority/important-soon',
105+
'priority/important-longterm',
106+
'epic',
107+
'lifecycle/frozen',
108+
];
109+
110+
// The sweep is the only place this workflow writes in bulk, and
111+
// GitHub's secondary rate limits allow 80 content-generating
112+
// requests a minute and 500 an hour. This client has no
113+
// throttling plugin and no retries, and a throttled write
114+
// answers 403, which is exempt from retry anyway: an unpaced
115+
// sweep does not slow down when it is throttled, it fails the
116+
// job and leaves the backlog half labeled. Pace the writes
117+
// under the per-minute ceiling, and stop short of the hourly
118+
// one instead of walking into it. Whatever a run leaves behind
119+
// stays unlabeled, so the next one picks it up.
120+
//
121+
// Those are the secondary limits. GITHUB_TOKEN also carries a
122+
// primary one of 1000 requests an hour for the whole repository,
123+
// shared with every other workflow, and a run that reaches the
124+
// cap spends a large part of it. That is affordable at 05:53 and
125+
// only while the backlog drains, but it is the ceiling to lower
126+
// the cap against if this ever moves to a busier slot.
127+
const WRITE_INTERVAL_MS = 1000;
128+
const MAX_WRITES_PER_RUN = 400;
129+
const sleep = (ms) =>
130+
new Promise((resolve) => setTimeout(resolve, ms));
131+
132+
// Returns the label to add, or null when the item needs nothing.
133+
// An issue somebody prioritised, assigned, made an epic of, or
134+
// froze has already been looked at by a maintainer, so recording
135+
// that as `triage/accepted` matches reality better than asking
136+
// for a triage that has effectively happened.
137+
const classify = (issue) => {
138+
if (issue.pull_request) return null; // listForRepo returns PRs too
139+
const labels = (issue.labels || []).map((l) =>
140+
typeof l === 'string' ? l : l.name
141+
);
142+
if (labels.some((n) => n.startsWith('triage/'))) return null;
143+
const looked = // somebody has already looked at this one
144+
(issue.assignees || []).length > 0 ||
145+
labels.some((n) => ACCEPTED_SIGNAL_LABELS.includes(n));
146+
return looked ? ACCEPTED : NEEDS_TRIAGE;
147+
};
148+
149+
const apply = async (issue, label, dryRun) => {
150+
if (dryRun) {
151+
core.info(`[dry-run] #${issue.number}: would add ${label}`);
152+
return;
153+
}
154+
await github.rest.issues.addLabels({
155+
owner: context.repo.owner,
156+
repo: context.repo.repo,
157+
issue_number: issue.number,
158+
labels: [label],
159+
});
160+
core.info(`#${issue.number}: added ${label}`);
161+
};
162+
163+
if (context.eventName === 'issues') {
164+
const issue = context.payload.issue;
165+
const label = classify(issue);
166+
if (label) {
167+
await apply(issue, label, false);
168+
} else {
169+
core.info(`#${issue.number}: already triaged, nothing to do`);
170+
}
171+
return;
172+
}
173+
174+
// schedule / workflow_dispatch: sweep every open issue.
175+
// workflow_dispatch inputs arrive as strings in the webhook
176+
// payload even when declared `type: boolean`.
177+
//
178+
// Keyed on the event rather than on the input alone, so both
179+
// unknown cases fail safe: `schedule` is the one event that
180+
// writes without being asked, so any trigger added here later
181+
// starts out dry rather than writing on its first run, and a
182+
// dispatch sending anything but the literal `false` gets a dry
183+
// run. The declared default above is what a human sees in the
184+
// form; this does not depend on it reaching the payload.
185+
const dryRun =
186+
context.eventName !== 'schedule' &&
187+
String(context.payload.inputs?.dry_run ?? 'true') !== 'false';
188+
189+
const issues = await github.paginate(github.rest.issues.listForRepo, {
190+
owner: context.repo.owner,
191+
repo: context.repo.repo,
192+
state: 'open',
193+
per_page: 100,
194+
});
195+
196+
let accepted = 0;
197+
let needsTriage = 0;
198+
let skipped = 0;
199+
let deferred = 0;
200+
let failed = 0;
201+
let refused = 0;
202+
let writes = 0;
203+
for (const issue of issues) {
204+
const label = classify(issue);
205+
if (!label) {
206+
skipped += 1;
207+
continue;
208+
}
209+
// The cap counts in dry-run too, so a dry run predicts what a
210+
// real one would do rather than a longer list nobody gets.
211+
if (writes >= MAX_WRITES_PER_RUN) {
212+
deferred += 1;
213+
continue;
214+
}
215+
if (writes > 0 && !dryRun) {
216+
await sleep(WRITE_INTERVAL_MS);
217+
}
218+
writes += 1;
219+
try {
220+
await apply(issue, label, dryRun);
221+
} catch (err) {
222+
// 403 is what a secondary rate limit and a token without
223+
// `issues: write` both answer; 429 is the other throttle
224+
// shape. Those two are the only statuses that stop the run.
225+
// Every other status carries on, and that is the whole of
226+
// the else: the 404 or 410 of an issue transferred or
227+
// deleted mid-sweep, but a 422 or a 5xx just the same,
228+
// because a failure about one issue says nothing about the
229+
// next one and the next sweep retries it. A refusal is the
230+
// opposite kind: it applies to every write left in the run,
231+
// and repeating a rejected request up to the cap is what
232+
// gets an integration banned, so stop at the first one. The
233+
// remainder stays unlabeled and the next run takes it,
234+
// exactly like the write cap's own deferral.
235+
if (err.status === 403 || err.status === 429) {
236+
refused += 1;
237+
core.warning(
238+
`#${issue.number}: ${label} refused ` +
239+
`(${err.status}); stopping the sweep`
240+
);
241+
break;
242+
}
243+
// One issue that cannot be labeled, because it was moved or
244+
// deleted mid-sweep, must not cost the run the rest of the
245+
// backlog and the summary with it. It stays unlabeled, so
246+
// the next run picks it up like any other. Counted below the
247+
// refusal branch rather than above it, so that one refusal
248+
// reports as one problem instead of also inflating the
249+
// failed count it is already reported by.
250+
failed += 1;
251+
core.warning(`#${issue.number}: ${label} failed: ${err.message}`);
252+
continue;
253+
}
254+
if (label === ACCEPTED) {
255+
accepted += 1;
256+
} else {
257+
needsTriage += 1;
258+
}
259+
}
260+
core.info(
261+
`Sweep done${dryRun ? ' (dry-run)' : ''}: ` +
262+
`${accepted} ${ACCEPTED}, ${needsTriage} ${NEEDS_TRIAGE}, ` +
263+
`${skipped} already triaged or PRs`
264+
);
265+
if (deferred > 0) {
266+
core.info(
267+
`${deferred} left for the next run: this one reached its ` +
268+
`write cap of ${MAX_WRITES_PER_RUN}`
269+
);
270+
}
271+
if (failed > 0) {
272+
core.info(`${failed} write(s) failed, left for the next run`);
273+
}
274+
// An issue that vanished mid-sweep is benign however many times
275+
// it happens, and the next run simply does not see it. A refusal
276+
// is not benign even once: it means the token lost `issues:
277+
// write`, or the pacing above stopped being enough and the sweep
278+
// is into the secondary rate limit. Discriminating on the status
279+
// rather than on a count keeps the alarm live once the backlog is
280+
// drained and a run writes to a handful of issues at most, which
281+
// is the regime this workflow spends almost all of its life in.
282+
if (refused > 0) {
283+
core.setFailed(
284+
`A label write was refused after ${writes} attempted, so ` +
285+
`the sweep stopped early; check the job's issues: write ` +
286+
`permission and the log for rate-limit responses`
287+
);
288+
}

0 commit comments

Comments
 (0)