forge-media
Image, transcription, speech, and video generation for the Forge SDK
Image, transcription, speech, and video generation for the Forge SDK
Package contract
| Field | Value |
|---|---|
| Language | rust |
| Source version | 0.2.0 |
| Manifest | forge-rs/crates/forge-media/Cargo.toml |
| Source files | 6 |
| Evidence | Source reference; registry publication and runtime conformance are separate checks |
Import boundary
use forge_media;Use a source checkout or your verified private registry. Manifest coordinates identify the package; they do not establish that a public registry release exists.
Crate boundary
The following entries are taken from src/lib.rs. Feature conditions in the exact source still apply.
pub mod error;
pub mod image;
pub mod speech;
pub mod transcription;
pub mod video;
pub use error::ForgeMediaError;
pub mod prelude;
pub use crate::error::ForgeMediaError;
pub use crate::image::{generate_image, ImageFormat, ImageOptions, ImageProvider, ImageResult};
pub use crate::speech::{speak, AudioFormat, SpeechOptions, SpeechProvider, SpeechResult};
pub use crate::transcription::{
transcribe, TranscriptionOptions, TranscriptionProvider, TranscriptionResult,
TranscriptionSegment,
};
pub use crate::video::{generate_video, VideoFormat, VideoOptions, VideoProvider, VideoResult};Source reference
Download package reference JSON. Each original source file and generated declaration artifact has its own SHA-256 digest. Function bodies and constant values are omitted from downloads. These are source declaration inventories, not compiler-resolved rustdoc, TypeDoc, DocC, or Dokka output. Private modules can contain public declarations that are not reachable through the package boundary; consult the entry point before importing.
error.rs
Read declaration text · 1 declaration entries
#[derive(Debug, Error)]
pub enum ForgeMediaError {
/// Image generation failed.
///
/// Returned when the underlying `ImageProvider::generate_image()` call
/// fails, including provider-side content filters, rate limits, and
/// invalid prompt errors.
#[error("image generation failed for model '{model}': {reason}")]
GenerationFailed {
/// The model identifier (e.g., "dall-e-3").
model: String,
/// The error message from the provider.
reason: String,
},
/// Audio transcription failed.
///
/// Returned when the underlying `TranscriptionProvider::transcribe()` call
/// fails, including invalid audio data, unsupported codecs, and provider
/// errors.
#[error("transcription failed for model '{model}': {reason}")]
TranscriptionFailed {
/// The model identifier (e.g., "whisper-1").
model: String,
/// The error message from the provider.
reason: String,
},
/// Speech synthesis failed.
///
/// Returned when the underlying `SpeechProvider::speak()` call fails,
/// including invalid voice identifiers, excessively long input, and
/// provider errors.
#[error("speech synthesis failed for model '{model}': {reason}")]
SpeechFailed {
/// The model identifier (e.g., "tts-1").
model: String,
/// The error message from the provider.
reason: String,
},
/// Video generation failed.
///
/// Returned when the underlying `VideoProvider::generate_video()` call
/// fails, including content filters, invalid dimensions, and provider
/// errors.
#[error("video generation failed for model '{model}': {reason}")]
VideoFailed {
/// The model identifier (e.g., "sora-1").
model: String,
/// The error message from the provider.
reason: String,
},
/// The requested media format is not supported.
///
/// Returned when a format is requested that the provider does not support,
/// or when input data uses an unrecognized encoding.
#[error("unsupported format '{format}': {reason}")]
UnsupportedFormat {
/// The format that was requested (e.g., "tiff", "aac").
format: String,
/// Why the format is not supported.
reason: String,
},
/// A `forge-core` error occurred during a media operation.
#[error("core error: {0}")]
Core(#[from] forge_core::error::ForgeError),
}image.rs
Read declaration text · 7 declaration entries
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum ImageFormat {
/// PNG (Portable Network Graphics) — lossless compression.
#[default]
Png,
/// JPEG — lossy compression, smaller file sizes.
Jpeg,
/// WebP — modern format with both lossy and lossless modes.
Webp,
}
pub fn mime_type(&self) -> &'static str;
pub fn extension(&self) -> &'static str;
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageOptions {
/// Desired width in pixels.
pub width: u32,
/// Desired height in pixels.
pub height: u32,
/// Output image format.
pub format: ImageFormat,
/// Output quality (0-100). Applicable to lossy formats like JPEG and WebP.
/// `None` means the provider's default quality.
pub quality: Option<u8>,
/// Style hint for the generation model (e.g., "photorealistic", "watercolor").
/// `None` means the provider's default style.
pub style: Option<String>
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageResult {
/// The raw image bytes in the format specified by `mime_type`.
pub data: Vec<u8>,
/// The MIME type of the image data (e.g., "image/png").
pub mime_type: String,
/// The width of the generated image in pixels.
pub width: u32,
/// The height of the generated image in pixels.
pub height: u32,
/// The model that generated this image (e.g., "dall-e-3").
pub model: String
}
#[async_trait]
pub trait ImageProvider: Send + Sync {
/// Returns the model identifier (e.g., "dall-e-3", "stable-diffusion-xl").
fn model_id(&self) -> &str;
/// Returns the provider name (e.g., "openai", "stability").
fn provider_name(&self) -> &str;
/// Generates an image from a text prompt.
///
/// # Arguments
///
/// * `prompt` - The text description of the image to generate.
/// * `options` - Configuration for output dimensions, format, quality, and style.
///
/// # Returns
///
/// An [`ImageResult`] containing the generated image bytes and metadata.
///
/// # Errors
///
/// * [`ForgeMediaError::GenerationFailed`] -- if the provider returns an error.
/// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
async fn generate_image(
&self,
prompt: &str,
options: &ImageOptions,
) -> Result<ImageResult, ForgeMediaError>;
}
pub async fn generate_image(
provider: &dyn ImageProvider,
prompt: &str,
options: &ImageOptions,
) -> Result<ImageResult, ForgeMediaError>;lib.rs
Read declaration text · 12 declaration entries
pub mod error;
pub mod image;
pub mod speech;
pub mod transcription;
pub mod video;
pub use error::ForgeMediaError;
pub mod prelude;
pub use crate::error::ForgeMediaError;
pub use crate::image::{generate_image, ImageFormat, ImageOptions, ImageProvider, ImageResult};
pub use crate::speech::{speak, AudioFormat, SpeechOptions, SpeechProvider, SpeechResult};
pub use crate::transcription::{
transcribe, TranscriptionOptions, TranscriptionProvider, TranscriptionResult,
TranscriptionSegment,
};
pub use crate::video::{generate_video, VideoFormat, VideoOptions, VideoProvider, VideoResult};speech.rs
Read declaration text · 7 declaration entries
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum AudioFormat {
/// MP3 — widely supported lossy audio format.
#[default]
Mp3,
/// WAV — uncompressed PCM audio.
Wav,
/// OGG — open container format, typically with Vorbis or Opus codec.
Ogg,
/// FLAC — lossless audio compression.
Flac,
}
pub fn mime_type(&self) -> &'static str;
pub fn extension(&self) -> &'static str;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct SpeechOptions {
/// The voice identifier to use (e.g., "alloy", "echo", "nova").
/// `None` means the provider's default voice.
pub voice: Option<String>,
/// The playback speed multiplier (e.g., 0.5 for half speed, 2.0 for double speed).
/// `None` means the provider's default speed (typically 1.0).
pub speed: Option<f64>,
/// Output audio format.
pub format: AudioFormat
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct SpeechResult {
/// The raw audio bytes in the format specified by `mime_type`.
pub audio: Vec<u8>,
/// The MIME type of the audio data (e.g., "audio/mpeg").
pub mime_type: String,
/// The duration of the generated audio in seconds, if available.
/// `None` if the provider does not report duration.
pub duration_seconds: Option<f64>
}
#[async_trait]
pub trait SpeechProvider: Send + Sync {
/// Returns the model identifier (e.g., "tts-1", "tts-1-hd").
fn model_id(&self) -> &str;
/// Returns the provider name (e.g., "openai", "elevenlabs").
fn provider_name(&self) -> &str;
/// Synthesizes speech from text.
///
/// # Arguments
///
/// * `text` - The text to convert to speech.
/// * `options` - Configuration for voice, speed, and output format.
///
/// # Returns
///
/// A [`SpeechResult`] containing the generated audio bytes and metadata.
///
/// # Errors
///
/// * [`ForgeMediaError::SpeechFailed`] -- if the provider returns an error.
/// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
async fn speak(
&self,
text: &str,
options: &SpeechOptions,
) -> Result<SpeechResult, ForgeMediaError>;
}
pub async fn speak(
provider: &dyn SpeechProvider,
text: &str,
options: &SpeechOptions,
) -> Result<SpeechResult, ForgeMediaError>;transcription.rs
Read declaration text · 6 declaration entries
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct TranscriptionSegment {
/// The start time of this segment in seconds from the beginning of the audio.
pub start: f64,
/// The end time of this segment in seconds from the beginning of the audio.
pub end: f64,
/// The transcribed text for this segment.
pub text: String
}
pub fn duration(&self) -> f64;
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct TranscriptionOptions {
/// An ISO 639-1 language code hint (e.g., "en", "fr", "de").
/// `None` means the provider will auto-detect the language.
pub language: Option<String>,
/// An optional prompt to guide the transcription model. Useful for providing
/// context about domain-specific terminology or expected content.
/// `None` means no guidance prompt.
pub prompt: Option<String>
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TranscriptionResult {
/// The full transcribed text.
pub text: String,
/// The detected language as an ISO 639-1 code, if available.
pub language: Option<String>,
/// The total duration of the audio in seconds, if available.
pub duration_seconds: Option<f64>,
/// Time-aligned transcription segments. May be empty if the provider
/// does not support segmented output.
pub segments: Vec<TranscriptionSegment>
}
#[async_trait]
pub trait TranscriptionProvider: Send + Sync {
/// Returns the model identifier (e.g., "whisper-1", "chirp-v2").
fn model_id(&self) -> &str;
/// Returns the provider name (e.g., "openai", "google").
fn provider_name(&self) -> &str;
/// Transcribes audio data to text.
///
/// # Arguments
///
/// * `audio` - Raw audio bytes. The provider determines acceptable formats
/// (e.g., WAV, MP3, FLAC, OGG).
/// * `options` - Configuration for language hint and guidance prompt.
///
/// # Returns
///
/// A [`TranscriptionResult`] containing the transcribed text and metadata.
///
/// # Errors
///
/// * [`ForgeMediaError::TranscriptionFailed`] -- if the provider returns an error.
/// * [`ForgeMediaError::UnsupportedFormat`] -- if the audio format is not supported.
async fn transcribe(
&self,
audio: &[u8],
options: &TranscriptionOptions,
) -> Result<TranscriptionResult, ForgeMediaError>;
}
pub async fn transcribe(
provider: &dyn TranscriptionProvider,
audio: &[u8],
options: &TranscriptionOptions,
) -> Result<TranscriptionResult, ForgeMediaError>;video.rs
Read declaration text · 7 declaration entries
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum VideoFormat {
/// MP4 — widely supported video container format (typically H.264/AAC).
#[default]
Mp4,
/// WebM — open video format (typically VP8/VP9/Opus).
Webm,
}
pub fn mime_type(&self) -> &'static str;
pub fn extension(&self) -> &'static str;
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct VideoOptions {
/// Desired width in pixels.
pub width: u32,
/// Desired height in pixels.
pub height: u32,
/// Desired video duration in seconds.
pub duration_seconds: f64,
/// Output video format.
pub format: VideoFormat
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct VideoResult {
/// The raw video bytes in the format specified by `mime_type`.
pub data: Vec<u8>,
/// The MIME type of the video data (e.g., "video/mp4").
pub mime_type: String,
/// The duration of the generated video in seconds.
pub duration_seconds: f64,
/// The width of the generated video in pixels.
pub width: u32,
/// The height of the generated video in pixels.
pub height: u32
}
#[async_trait]
pub trait VideoProvider: Send + Sync {
/// Returns the model identifier (e.g., "sora-1", "runway-gen2").
fn model_id(&self) -> &str;
/// Returns the provider name (e.g., "openai", "runway").
fn provider_name(&self) -> &str;
/// Generates a video from a text prompt.
///
/// # Arguments
///
/// * `prompt` - The text description of the video to generate.
/// * `options` - Configuration for output dimensions, duration, and format.
///
/// # Returns
///
/// A [`VideoResult`] containing the generated video bytes and metadata.
///
/// # Errors
///
/// * [`ForgeMediaError::VideoFailed`] -- if the provider returns an error.
/// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
async fn generate_video(
&self,
prompt: &str,
options: &VideoOptions,
) -> Result<VideoResult, ForgeMediaError>;
}
pub async fn generate_video(
provider: &dyn VideoProvider,
prompt: &str,
options: &VideoOptions,
) -> Result<VideoResult, ForgeMediaError>;