String -> Kokoro Voice
Turn 'Heart' or 'Roger' into a real Kokoro voice
- voice
- pitch
StringToKokoroVoice is a tiny lookup node with an outsized job: it turns the short, human-readable character names in your story JSON into the concrete voice ID and pitch shift that KokoroSceneAudio actually needs. In the Story JSON Nodes pack (stepan-bogatorjov/comfy_scenes_json_node), scenes pick a narrator by a friendly name like "Heart", "Fenrir" or "Child" - this node is what makes those names mean something.
Kokoro (hexgrad's 82M model) has no notion of "Heart". It has 28 voice IDs with prefixes that encode dialect and gender - af_ American female, am_ American male, bf_/bm_ British. This node maps the friendly name to one of those IDs, and where the name describes something Kokoro doesn't have a dedicated voice for (a child, say), it also emits a pitch shift in semitones to get there. "Child" resolves to the light af_sky voice pitched up 6 semitones - a convincing kid timbre from a voice that wasn't built for it.
How it works
One deterministic resolve() function, no randomness. A curated table maps the good names to (voice, pitch): "Heart" → af_heart +0, "Fenrir" → am_fenrir +0, "Fable" → bm_fable +0, the child variants with their semitone shifts. A raw Kokoro ID passes straight through. Legacy ElevenLabs-style names ("Roger", "Laura", "Will") map to the nearest Kokoro voice for backwards compatibility. Anything unknown falls back to am_michael. Curated names also match case-insensitively as a convenience.
Inputs and outputs
- name - the string from your story JSON (default "Heart").
- voice (STRING) → wire into KokoroSceneAudio's
voice_name. - pitch (FLOAT) → wire into KokoroSceneAudio's
pitch.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/stepan-bogatorjov/comfy_scenes_json_node
Restart ComfyUI (or install via Manager under "Story JSON Nodes"). No model, no network - it's a pure lookup.
Troubleshooting
The subtle trap: KokoroSceneAudio's _resolve_voice uses the mapped voice ID but ignores the mapper's pitch - the pitch arrives through the node's own pitch input, which is fed by this node's pitch output. So if you wire voice but forget to wire pitch, adult voices work fine but "Child" and "ChildWarm" come out at adult pitch. Wire both sockets. And note it's just a string mapper: nothing here downloads or runs a model, so if you see silence downstream, the problem is in KokoroSceneAudio's setup, not this node.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| name | STRING | Heart | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| voice | STRING | — |
| pitch | FLOAT | — |