profiler
The profiler model — the single source of truth that turns the engine's flat profiling surface (`getProfilingDataTable`, `getLuauVmDataTable`, `profiler.stopCapture`) into the honest, self-accounting hierarchy every profiler tool and the System Tools Profiler tab reads from.
profiler
The profiler model — the single source of truth that turns the engine's flat profiling surface (getProfilingDataTable, getLuauVmDataTable, profiler.stopCapture) into the honest, self-accounting hierarchy every profiler tool and the System Tools Profiler tab reads from.
Location: src/lua/lib/modules/profiler.module
Required as require("modules.profiler").
What it composes
- A nested tree where every node carries
total_ms,self_ms(own time excluding children),calls, andpctof the frame. - An explicit
present / idlenode — the frame time no schedule covers (the frame-limiter / vsync wait, GPU present, event loop) — so nothing vanishes. - Per-component
update(dt)attribution grafted underlua_update.vm_call, so a heavy script names the entity + component an author can open and fix. - Per-component coroutine-resume attribution under
task_scheduler, so a busy scheduler names the component whose tasks it resumed. - Windowed aggregation over a capture: avg / p50 / p90 / p99 / max frame time, the worst single frame, and spike counting.
The model + format functions are pure — the tools and the UI produce identical numbers from identical inputs. See the profiler toolbox for the tools that read this model.
Interface
What this asset declares: the schema it conforms to, what it exposes, and the rendered structured payload.
conforms to
zero/source-extract/v2Profiler model The single source of truth for turning the engine's raw profiling data into the honest, self-accounting hierarchy every profiler tool and the System Tools Profiler tab reads from. The raw engine surface is deliberately flat: `getProfilingDataTable()` gives per-schedule and per-system wall-clock keyed by dotted name (`lua_update`, `lua_update.vm_call`, ...); `getLuauVmDataTable()` gives per-component-instance `update(dt)` cost; `profiler.stopCapture()` gives a window of per-frame samples. This module composes those into: * a nested tree where every node carries `total_ms`, `self_ms` (own time excluding children), `calls`, and `pct` of the frame, * an explicit `present / idle` node = frame time no schedule covers (the frame-limiter / vsync wait, GPU present, event loop), * per-component `update(dt)` attribution grafted under `lua_update.vm_call` so a heavy script names the entity + component an author can go fix, * windowed aggregation over a capture (avg / p50 / p90 / p99 / max frame time, the worst single frame, spike counting). Nothing here reads engine state as a side effect except the `live*` / `readCapture` acquisition helpers; the model + format functions are pure so the tools and the UI produce identical numbers from identical inputs.
engineGlobal(name: ?) → void
Acquisition (the only functions that read live engine state) Resolve an engine global installed before `_G` was sealed.
| arg | type | description |
|---|---|---|
| name | ? |
liveProfiling( ) → any
The live `getProfilingDataTable()` snapshot, or nil if unavailable.
liveScripts(minMs: number?) → void
Live per-component `update(dt)` rows from `getLuauVmDataTable()`, normalized to `{ label, component_type, entity_id, source_path, total_ms, avg_ms, max_ms, calls }` and sorted by `total_ms` descending. Rows below `minMs` (default 0) are dropped.
| arg | type | description |
|---|---|---|
| minMs | number? |
liveTasks(minMs: number?) → void
Live per-component coroutine-resume rows from `getLuauVmDataTable().task_details`, normalized to the same leaf shape as `liveScripts` (`{ label, name, entity_id, source_path, total_ms, calls, kind = "task" }`) and sorted by cost descending. These attribute the `task_scheduler` cost to the component whose tasks were resumed. Rows below `minMs` (default 0) are dropped.
| arg | type | description |
|---|---|---|
| minMs | number? |
rollupByType(rows: { any }) → void
Roll per-instance rows (from `liveScripts`/`liveTasks`) up to per-component type: sum `total_ms`/`calls`, count instances, keep the peak. A type with a single instance keeps its `@ entity` label; many instances collapse to `Type xN`. Sorted by cost. This is what keeps the view bounded when a world has thousands of instances of a handful of component types.
| arg | type | description |
|---|---|---|
| rows | { any } |
capLeaves(rows: { any }, n: number?, noun: string?) → void
Keep the top `n` rows; fold the rest into one `(k more ...)` row so a grafted list stays bounded while its summed total (and the parent's self math) is preserved. `noun` labels the overflow row.
| arg | type | description |
|---|---|---|
| rows | { any } | |
| n | number? | |
| noun | string? |
liveFrame(opts: any?) → any
Assemble the live attributed frame tree with component-type rollup + top-N capping on the script and task grafts, so the tree is bounded no matter how many instances a world has. `opts = { minMs = 0.1, topLeaves = 12 }`.
| arg | type | description |
|---|---|---|
| opts | any? |
liveVm( ) → any
Full VM introspection table (memory + scripts) from `getLuauVmDataTable()`.
readCapture(label: string?) → any
Parse a capture (a `profiler.stopCapture()`/`lastCapture()` JSON string, or the VFS file a named capture was written to). `label` nil => last capture.
| arg | type | description |
|---|---|---|
| label | string? |
currentCallerId( ) → string
The caller this call is attributed to: the same key the profiler records a frame's queued-Luau cost under. Nil when the transport named no caller.
callerCosts(frames: { any }, selfCaller: string?, frameCount: number?) → void
Roll each caller's per-frame cost up across a capture's frames. Returns `{ { id, total_ms, avg_ms, max_ms, frames, is_self, is_engine } }`, heaviest first, over `frameCount` frames, with `selfCaller` marked.
| arg | type | description |
|---|---|---|
| frames | { any } | |
| selfCaller | string? | |
| frameCount | number? |
foreignCallerLoad(frames: { any }, selfCaller: string?) → void
How many of `frames` ran queued Luau belonging to a caller other than `selfCaller`, and what that cost per frame on average. Engine-level work belongs to no session, so it counts as neither.
| arg | type | description |
|---|---|---|
| frames | { any } | |
| selfCaller | string? |
describeRecordingHold(hold: any, selfCaller: string?) → string
Say who holds the one windowed recording. `hold` is the decoded `profiler.recordingHold()` table; `selfCaller` the reading caller's own id. Returns `"'alpha', held by <who>, running for 12.3s"`.
| arg | type | description |
|---|---|---|
| hold | any | |
| selfCaller | string? |
holdIsMine(hold: any, selfCaller: string?) → boolean
Whether `hold` names a caller, and whether that caller is `selfCaller`.
| arg | type | description |
|---|---|---|
| hold | any | |
| selfCaller | string? |
formatRecordingHeld(hold: any, selfCaller: string?, wanted: string) → string
Why a `record start <wanted>` was refused: which recording holds the one engine-wide slot, who holds it, and what the blocked caller can do that does not end somebody else's measurement.
| arg | type | description |
|---|---|---|
| hold | any | |
| selfCaller | string? | |
| wanted | string |
formatRecordingStatus(hold: any, selfCaller: string?) → string
What `record status` reports: the holder when one exists, or how to start. A holder reading its own recording is told how to stop it; a caller reading somebody else's is told the same thing the refusal tells it, because the stop command ends a measurement that is not its to end.
| arg | type | description |
|---|---|---|
| hold | any | |
| selfCaller | string? |
formatRingWindowShortfall(heldSeconds: number?, askedSeconds: number?) → string
How much history a retro window actually covers against how much it asked for. The ring is a bounded, always-overwriting buffer, so a window can hold less than the seconds requested: the ring retains fewer, or the engine produced fewer frames in that span than a healthy one would. Returns nil when the window covers what was asked.
| arg | type | description |
|---|---|---|
| heldSeconds | number? | |
| askedSeconds | number? |
profDataFromFlat(flat: ?, dtMs: ?) → void
Tree assembly (pure) Reshape a flat `{ dottedKey = ms }` map (as a capture frame stores it) into the `{ frame, schedules, systems }` structure `getProfilingDataTable()` returns, so one `buildTree` serves both live snapshots and captured frames.
| arg | type | description |
|---|---|---|
| flat | ? | |
| dtMs | ? |
nestSystems(sysMap: ?) → void
Nest a flat map of dotted keys into tree nodes. Each node's `total` is its own measured `last_ms`; a child's time is a slice of its parent's.
| arg | type | description |
|---|---|---|
| sysMap | ? |
child(arr: ?, map: ?, seg: ?) → void
| arg | type | description |
|---|---|---|
| arr | ? | |
| map | ? | |
| seg | ? |
isLeafKind(kind: ?) → void
A grafted leaf (a script or a task) is terminal — never recursed into.
| arg | type | description |
|---|---|---|
| kind | ? |
finalizeNode(node: ?, dtMs: ?, graft: ?) → void
Recursively finalize a node: graft any attributed leaves (scripts under vm_call, tasks under task_scheduler), sort children, compute self_ms (total - sum children) and pct of the frame. `graft` maps a node's dotted key to the list of leaf rows that belong under it.
| arg | type | description |
|---|---|---|
| node | ? | |
| dtMs | ? | |
| graft | ? |
attributeResidual(residualMs: number, cpuMs: number?, gpuMs: number?, gpuSupported: boolean?) → any
Explain a frame's unattributed wait using how busy the GPU was during it. `present / idle` is a residual — whatever the CPU schedules did not account for. Two very different frames land in it: one where the GPU is still working and present blocks behind it, and one where the CPU has nothing to do and waits on the display's pace. They demand opposite responses, so this names which one it is and reports the GPU number it decided from. The GPU runs alongside the CPU, so the GPU work the CPU's own schedules already spanned is hidden behind them however large it is — only the part that runs past them (`gpuMs - cpuMs`) can be what the frame waits on. That overhang is what the verdict is taken from, capped at the wait it has to fit inside. `cpuMs` is the frame's summed schedule cost and `gpuMs` the GPU work in the frame (`profiler.gpuFrame().total_ms`). `gpuSupported` false means the backend has no timestamp queries, which is reported as `unknown` rather than guessed at. Returns `{ kind, gpuMs, overhangMs, share, note }`, where `kind` is one of `"gpu-bound"`, `"paced"`, `"unknown"` or `"none"`, and `share` is the fraction of the wait the overhang covers. `M.RESIDUAL_LABEL[kind]` names the residual row for each verdict.
| arg | type | description |
|---|---|---|
| residualMs | number | |
| cpuMs | number? | |
| gpuMs | number? | |
| gpuSupported | boolean? |
buildTree(profData: any, scripts: { any }?, tasks: { any }?, gpu: any?) → any
Build the attributed frame model from a `getProfilingDataTable()`-shaped table plus (optional) normalized script rows (grafted under `vm_call`), task rows (grafted under `task_scheduler`) and the frame's GPU timing (`profiler.gpuFrame()`), which explains the residual. Returns `{ dt_ms, fps, present_idle_ms, roots = { Node }, scripts, tasks }` where each `Node = { name, full, total_ms, self_ms, calls, pct, kind, children }`.
| arg | type | description |
|---|---|---|
| profData | any | |
| scripts | { any }? | |
| tasks | { any }? | |
| gpu | any? |
alias(node: ?) → void
Present each node's `total` under a uniform `total_ms` alias too, so downstream readers don't have to know the internal field name.
| arg | type | description |
|---|---|---|
| node | ? |
hotspots(frame: any, n: number?) → void
Derived views (pure) Flatten a frame model into the top-`n` nodes by SELF time (the actual expensive leaves — a system's own unattributed cost, or a script). Each row is `{ name, self_ms, total_ms, pct, kind, path }`. `path` locates the node (its schedule / parent chain) so a flat hit is still traceable.
| arg | type | description |
|---|---|---|
| frame | any | |
| n | number? |
walk(node: ?, path: ?) → void
| arg | type | description |
|---|---|---|
| node | ? | |
| path | ? |
percentile(values: ?, p: ?) → void
Percentile (0..100) of a numeric array using nearest-rank; array is copied+sorted.
| arg | type | description |
|---|---|---|
| values | ? | |
| p | ? |
effectiveMs(frame: any, excludeAgent: boolean?) → number
Effective frame time: wall-clock minus agent-injected cost (when `excludeAgent`). This is the frame time the actual game would have had without the agent's call.
| arg | type | description |
|---|---|---|
| frame | any | |
| excludeAgent | boolean? |
frameSignature(frame: any) → void
The dominant NON-agent hotspot (the system with the most SELF time) in a captured frame, plus the frame's built tree. This is the frame's "signature" — WHY it was slow — so frames that spiked for the same reason cluster together. Returns `{ name, path, self_ms }?, tree`.
| arg | type | description |
|---|---|---|
| frame | any |
walk(node: ?, path: ?) → void
| arg | type | description |
|---|---|---|
| node | ? | |
| path | ? |
domainMs(frame: any, membership: any) → number
Sum a frame's systems belonging to a domain (a `{ leaf = true }` membership set), across every schedule. `frame.systems` is a flat map of full dotted keys → this frame's `frame_ms`, so a domain that spans schedules (physics) rolls up into one honest number the per-schedule tree can't show.
| arg | type | description |
|---|---|---|
| frame | any | |
| membership | any |
aggregateCapture(capture: any, spikeMs: number?, opts: any?) → any
Aggregate a capture (per-frame samples) into windowed frame-time stats plus a DISTRIBUTION of what's slow. Rather than a single anecdotal worst frame (which over a long capture is likely a random hiccup), every spike frame is grouped by its dominant hotspot into `clusters` — so one capture answers "what's slow" completely. All statistics are on EFFECTIVE frame time (wall-clock minus the agent's own injected `execute` cost) unless `opts.excludeAgent == false`. `spikeMs` (default: 1.5x the median) marks a frame as a spike. Several agent sessions drive one engine, so the window's queued-Luau cost is split across the callers that asked for it: `callers` names each one's share and `foreign_frames` counts the frames that carried a caller other than the reader's own. `opts.selfCaller` is the reading caller's id, defaulting to the live one; pass `""` to read as nobody. Returns `{ label, source, frames, seconds, exclude_agent, agent_frames, dt = {avg,p50,p90,p99,max,min}, spike_ms, spikes, clusters = { { signature, path, count, avgMs, maxMs, share_pct, worstIndex, tree } }, worst = { index, dt_ms, tree }, typical = { tree }, callers = { { id, total_ms, avg_ms, max_ms, frames, is_self, is_engine } }, caller_self, foreign_frames, foreign_ms }`.
| arg | type | description |
|---|---|---|
| capture | any | |
| spikeMs | number? | |
| opts | any? |
ms(v: ?) → void
Formatting (pure text; the tools return this as stdout)
| arg | type | description |
|---|---|---|
| v | ? |
pct(v: ?) → void
| arg | type | description |
|---|---|---|
| v | ? |
formatTree(frame: any, opts: any?) → string
Render a frame model as a columnar tree: `self | total | calls | % | name`. `opts = { minMs = 0.1, maxDepth = 32, header = true }`.
| arg | type | description |
|---|---|---|
| frame | any | |
| opts | any? |
line(node: ?, depth: ?) → void
| arg | type | description |
|---|---|---|
| node | ? | |
| depth | ? |
formatHotspots(rows: { any }) → string
Render hotspots as a flat table: `self | % | kind | name (path)`.
| arg | type | description |
|---|---|---|
| rows | { any } |
formatScripts(rows: { any }) → string
Render script attribution as a table: `last | avg | max | calls | component @ entity`.
| arg | type | description |
|---|---|---|
| rows | { any } |
formatCallerCosts(agg: any) → string
Render the window's queued-Luau cost split across the callers that asked for it, so a reader on a shared engine tells its own cost from a sibling session's and knows which rows below carry work it did not ask for.
| arg | type | description |
|---|---|---|
| agg | any |
formatRecordReport(agg: any) → string
Render a windowed capture aggregate: headline stats, the spike-cluster distribution (grouped by dominant hotspot, so one capture shows the WHOLE picture, not one anecdotal worst frame), the top cluster's representative tree, and the typical frame. All times are effective (agent cost excluded).
| arg | type | description |
|---|---|---|
| agg | any |
formatTasks(rows: { any }) → string
Render coroutine-resume attribution (task rows) as a table: `resume | calls | component @ entity` (or `Type xN`).
| arg | type | description |
|---|---|---|
| rows | { any } |
formatMemory(vm: any, n: number?) → string
Render VM + per-script memory from a `getLuauVmDataTable()` table: the VM total + GC state, then the heaviest scripts by retained bytes.
| arg | type | description |
|---|---|---|
| vm | any | |
| n | number? |
formatCompare(labelA: string, aggA: any, labelB: string, aggB: any) → string
Diff two capture aggregates (from `aggregateCapture`): frametime deltas and the biggest per-system movers between the two recordings.
| arg | type | description |
|---|---|---|
| labelA | string | |
| aggA | any | |
| labelB | string | |
| aggB | any |
sysAvg(agg: ?) → void
| arg | type | description |
|---|---|---|
| agg | ? |
walk(n: ?) → void
| arg | type | description |
|---|---|---|
| n | ? |
Sub-parts
Everything contained inside this part. Assets are composite children (clickable cards). Files are leaf payloads. Expand any row to view its source.
Problems
Everything affecting this asset right now: its own problems, anything wrong inside it, and problems on its direct dependencies.
agent_score is exposed.+ quality × 0.35
+ performance × 0.25
± compat factor
Usability ratings
Did the part work as advertised when consumers tried to drop it in. Separate from upvotes: those are taste; this is "did it function".
Scoped to this part · feeds back into the world's score.