Compute and usage
There is nothing to set up. Anything heavy runs on Relay’s cloud and you pay for the time it runs; anything light runs in your browser for free. This page says which is which and where the numbers are.
What runs where
Section titled “What runs where”| Where | Costs | |
|---|---|---|
| Opening and viewing files | Your browser | Nothing |
| Process tools (blur, threshold, FFT…) | Your browser | Nothing |
| Notebook on CPU | A cloud machine | Compute time |
| Notebook on T4, A10G, A100 | That GPU | GPU hours, while the kernel is live |
| Pipelines and schedules | The server, or your browser if the pipeline needs a format only the browser reads | Compute time for server steps |
| Ray | Relay’s model provider | One AI prompt per question |
A notebook GPU is billed from the moment the kernel is Live until you press Stop. Take the big GPU for the training cell and drop back to CPU afterwards.
Where the numbers are
Section titled “Where the numbers are”Settings → Plan & usage → Usage has four cards:
- Allowances — what you have used of each limit on your plan: Storage, AI prompts (per day), GPU hours and Compute spend (per month).
- Ray activity — runs in the last day and month, tokens, and the estimated model cost; the list underneath is every run with its transcript. See Ray activity.
- Relay Drive — files, space left, and the size above which uploads go to the cloud automatically.
- Compute — active jobs and your concurrent-job limit.
Plans lists what each plan includes; Billing is where the card and the invoices live.
When a run is refused
Section titled “When a run is refused”A pipeline step or notebook kernel is checked before it starts. If it would go over your remaining GPU hours or budget, or you already have the maximum number of jobs running, it is refused with the reason and nothing is charged. Wait for a job to finish, or upgrade under Plans.