# Skill: captions/quiet — "שקט" (Captions only)

שורה בתוך קופסה שקופה ומעוגלת, מה שעוד לא נאמר חצי שקוף

**Font:** Assistant 600 · **Animation:** highlight · **Price:** 61–125 credits per finished minute (≈ ₪2.7–₪5.5) — one render;
a silence cut before it is a second render (the top of the range).

> **Given a captions skill AND a hook skill?** They are one job, not two. Put `captions` (from the captions
> skill) and `hook` (from the hook skill) in the **same** `POST /v1/studio/edit` call — one render, the price
> of captions alone. Do not render twice, and never send a video that already has burned-in captions back with
> `captions` again (it gets two sets). The ready-made combination:
> `https://apihub.co.il/skills/captions-quiet+hook-<effect>.md`

Base URL: `https://apihub.co.il` · Auth: `Authorization: Bearer <API key>` · Docs: `/docs`
Send the same `X-Tag: <job name>` header on every call of one video; `GET /v1/usage?tag=<job name>` then
returns exactly what it cost. Times in request bodies are **seconds on the edited (output) timeline**.
Unknown fields are rejected with the name of the right field — read the error instead of guessing.
**Files are kept 24 hours** (renders, uploads, generated clips/images/music) — download the finished video.

---

## 1. (Optional) cut the pauses first

```jsonc
POST /v1/studio/remove-silence     // x-wait-ms: 110000
{ "video_url": "<source>", "language_code": "he",
  "min_pause_ms": 350, "padding_ms": 70, "remove_fillers": true,
  "output": { "resolution": "source", "aspect_ratio": "original", "quality": "high" } }
```

Use the returned `url` below. Skip this step to keep the original pacing (and half the price).

## 2. Render the captions

```jsonc
POST /v1/studio/edit               // x-wait-ms: 110000
{
  "video_url": "<source or cut url>",
  "language_code": "he",
  "captions": { "style": "quiet" },
  "audio": { "normalize": true },
  "output": { "aspect_ratio": "9:16", "resolution": "1080", "fit": "crop", "quality": "high" }
}
```

Send `x-wait-ms: 110000` to wait for the render in the same call; if the response is still
`processing`, poll `GET /v1/executions/{id}` until `status` is `completed` and read `output.url`.

The transcription happens inside the call (AssemblyAI, word-level timing, 99 languages incl. Hebrew).
If you already have a transcript, pass `"transcript_id": "<id>"` instead of `language_code` and nothing is
transcribed twice.

## 3. What the preset is, exactly

`"style": "quiet"` expands to this. Send `"style": { "preset": "quiet", ...overrides }` to change any field —
the preset is only a starting point. Height: `position` + `margin`, or `x`/`y` in output pixels with
`anchor`. Size: `size` (% of the frame height). Type: `font`, `weight`, `letter_spacing`, `uppercase`.
Colours: `color`, `active_color`, `active_background`, `outline_color`, `shadow_color`, `background`.
Rhythm: `max_words`, `max_lines`, `max_chars_per_line`, `animation`:

```json
{
  "font": "Assistant",
  "weight": 600,
  "italic": false,
  "size": 5.2,
  "color": "#FFFFFF",
  "active_color": "#FFFFFF",
  "active_background": null,
  "outline_color": "#000000",
  "outline": 0,
  "shadow": 3,
  "shadow_color": "#00000099",
  "background": "#FFFFFF2E",
  "uppercase": false,
  "letter_spacing": 0,
  "position": "bottom",
  "margin": 14,
  "side_margin": 6,
  "max_chars_per_line": 30,
  "max_lines": 2,
  "max_words": 8,
  "animation": "highlight",
  "glow": 0,
  "hollow": false,
  "keywords": [],
  "keyword_scale": 1,
  "keyword_background": null,
  "future_opacity": 0.5,
  "past_color": null,
  "active_rotate": 0,
  "active_underline": 0,
  "underline_color": "#FFD600",
  "word_background": null,
  "word_radius": 0,
  "background_radius": 16,
  "skew": 0
}
```

**The knobs that matter for this look**

- `background` — the bar behind the whole line (now `#FFFFFF2E`), `background_radius` rounds it
- `future_opacity` — words not yet spoken are drawn at 50%; 1 turns that off
- `size` (% of the frame height, now 5.2), `position` + `margin` (now bottom / 14), `max_words` (now 8) and `max_lines` (now 2) shape the rhythm

`GET /v1/studio/caption-styles` lists every preset with its full settings; `POST /v1/studio/captions/preview`
renders one frame (`{ video_url, style, at, text }`) so you can check a look for ≈ 1 credit before the full render.

## 4. Add a hook (same render, same price)

A headline in the first seconds is burned into the same pass — add `"hook": { "text": "<5–8 words>",
"duration": 3, "style": "punch" }` to the call above; do not render again. `GET /v1/studio/hook-styles` lists
the 12 effects; each has its own skill at `/skills/hook-<effect>.md`, and this style with any of them is
`/skills/captions-quiet+hook-<effect>.md`.

## Rules that keep captions readable on every platform

- **Safe zone.** Keep text inside the middle 80% of the frame: `side_margin` ≥ 6 handles the sides; `margin`
  ≥ 12 keeps the bottom clear of the platform UI (username, buttons, progress bar).
- **A line never leaves the frame** — a line too wide for the side margins is set smaller automatically. Do
  not shrink the size yourself, wrap instead (`max_chars_per_line`).
- **Hebrew has no capitals** — keep `uppercase: false`; it only breaks Latin words mixed into a sentence.
- **Fonts.** `GET /v1/studio/fonts` lists the bundled families and which scripts they draw. A font that
  cannot draw the caption's language is swapped and reported in `warnings`.
- The customer's transcript wins: to fix a word, `GET /v1/speech/transcripts/{id}/captions`, edit the text,
  and send the list back as `captions.captions` — timing is kept when the word count is unchanged.

## What it costs

Transcription ≈ 1 credit per minute of audio. The render ≈ 1 credit per second of output. A 60-second video
captioned straight from the source ≈ 61 credits (≈ ₪2.7); cut first and it is ≈ 125
(≈ ₪5.5). Nothing else is billed — this skill generates nothing.
