Forge documentation
Library referenceRust

forge-media

Image, transcription, speech, and video generation for the Forge SDK

Image, transcription, speech, and video generation for the Forge SDK

Package contract

FieldValue
Languagerust
Source version0.2.0
Manifestforge-rs/crates/forge-media/Cargo.toml
Source files6
EvidenceSource reference; registry publication and runtime conformance are separate checks

Import boundary

use forge_media;

Use a source checkout or your verified private registry. Manifest coordinates identify the package; they do not establish that a public registry release exists.

Crate boundary

The following entries are taken from src/lib.rs. Feature conditions in the exact source still apply.

pub mod error;

pub mod image;

pub mod speech;

pub mod transcription;

pub mod video;

pub use error::ForgeMediaError;

pub mod prelude;

pub use crate::error::ForgeMediaError;

pub use crate::image::{generate_image, ImageFormat, ImageOptions, ImageProvider, ImageResult};

pub use crate::speech::{speak, AudioFormat, SpeechOptions, SpeechProvider, SpeechResult};

pub use crate::transcription::{
        transcribe, TranscriptionOptions, TranscriptionProvider, TranscriptionResult,
        TranscriptionSegment,
    };

pub use crate::video::{generate_video, VideoFormat, VideoOptions, VideoProvider, VideoResult};

Source reference

Download package reference JSON. Each original source file and generated declaration artifact has its own SHA-256 digest. Function bodies and constant values are omitted from downloads. These are source declaration inventories, not compiler-resolved rustdoc, TypeDoc, DocC, or Dokka output. Private modules can contain public declarations that are not reachable through the package boundary; consult the entry point before importing.

error.rs

Read declaration text · 1 declaration entries

#[derive(Debug, Error)]
pub enum ForgeMediaError {
    /// Image generation failed.
    ///
    /// Returned when the underlying `ImageProvider::generate_image()` call
    /// fails, including provider-side content filters, rate limits, and
    /// invalid prompt errors.
    #[error("image generation failed for model '{model}': {reason}")]
    GenerationFailed {
        /// The model identifier (e.g., "dall-e-3").
        model: String,
        /// The error message from the provider.
        reason: String,
    },

    /// Audio transcription failed.
    ///
    /// Returned when the underlying `TranscriptionProvider::transcribe()` call
    /// fails, including invalid audio data, unsupported codecs, and provider
    /// errors.
    #[error("transcription failed for model '{model}': {reason}")]
    TranscriptionFailed {
        /// The model identifier (e.g., "whisper-1").
        model: String,
        /// The error message from the provider.
        reason: String,
    },

    /// Speech synthesis failed.
    ///
    /// Returned when the underlying `SpeechProvider::speak()` call fails,
    /// including invalid voice identifiers, excessively long input, and
    /// provider errors.
    #[error("speech synthesis failed for model '{model}': {reason}")]
    SpeechFailed {
        /// The model identifier (e.g., "tts-1").
        model: String,
        /// The error message from the provider.
        reason: String,
    },

    /// Video generation failed.
    ///
    /// Returned when the underlying `VideoProvider::generate_video()` call
    /// fails, including content filters, invalid dimensions, and provider
    /// errors.
    #[error("video generation failed for model '{model}': {reason}")]
    VideoFailed {
        /// The model identifier (e.g., "sora-1").
        model: String,
        /// The error message from the provider.
        reason: String,
    },

    /// The requested media format is not supported.
    ///
    /// Returned when a format is requested that the provider does not support,
    /// or when input data uses an unrecognized encoding.
    #[error("unsupported format '{format}': {reason}")]
    UnsupportedFormat {
        /// The format that was requested (e.g., "tiff", "aac").
        format: String,
        /// Why the format is not supported.
        reason: String,
    },

    /// A `forge-core` error occurred during a media operation.
    #[error("core error: {0}")]
    Core(#[from] forge_core::error::ForgeError),
}

image.rs

Read declaration text · 7 declaration entries

#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum ImageFormat {
    /// PNG (Portable Network Graphics) — lossless compression.
    #[default]
    Png,
    /// JPEG — lossy compression, smaller file sizes.
    Jpeg,
    /// WebP — modern format with both lossy and lossless modes.
    Webp,
}

pub fn mime_type(&self) -> &'static str;

pub fn extension(&self) -> &'static str;

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageOptions {
/// Desired width in pixels.

pub width: u32,
/// Desired height in pixels.

pub height: u32,
/// Output image format.

pub format: ImageFormat,
/// Output quality (0-100). Applicable to lossy formats like JPEG and WebP.

/// `None` means the provider's default quality.

pub quality: Option<u8>,
/// Style hint for the generation model (e.g., "photorealistic", "watercolor").

/// `None` means the provider's default style.

pub style: Option<String>
}

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct ImageResult {
/// The raw image bytes in the format specified by `mime_type`.

pub data: Vec<u8>,
/// The MIME type of the image data (e.g., "image/png").

pub mime_type: String,
/// The width of the generated image in pixels.

pub width: u32,
/// The height of the generated image in pixels.

pub height: u32,
/// The model that generated this image (e.g., "dall-e-3").

pub model: String
}

#[async_trait]
pub trait ImageProvider: Send + Sync {
    /// Returns the model identifier (e.g., "dall-e-3", "stable-diffusion-xl").
    fn model_id(&self) -> &str;

    /// Returns the provider name (e.g., "openai", "stability").
    fn provider_name(&self) -> &str;

    /// Generates an image from a text prompt.
    ///
    /// # Arguments
    ///
    /// * `prompt` - The text description of the image to generate.
    /// * `options` - Configuration for output dimensions, format, quality, and style.
    ///
    /// # Returns
    ///
    /// An [`ImageResult`] containing the generated image bytes and metadata.
    ///
    /// # Errors
    ///
    /// * [`ForgeMediaError::GenerationFailed`] -- if the provider returns an error.
    /// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
    async fn generate_image(
        &self,
        prompt: &str,
        options: &ImageOptions,
    ) -> Result<ImageResult, ForgeMediaError>;
}

pub async fn generate_image(
    provider: &dyn ImageProvider,
    prompt: &str,
    options: &ImageOptions,
) -> Result<ImageResult, ForgeMediaError>;

lib.rs

Read declaration text · 12 declaration entries

pub mod error;

pub mod image;

pub mod speech;

pub mod transcription;

pub mod video;

pub use error::ForgeMediaError;

pub mod prelude;

pub use crate::error::ForgeMediaError;

pub use crate::image::{generate_image, ImageFormat, ImageOptions, ImageProvider, ImageResult};

pub use crate::speech::{speak, AudioFormat, SpeechOptions, SpeechProvider, SpeechResult};

pub use crate::transcription::{
        transcribe, TranscriptionOptions, TranscriptionProvider, TranscriptionResult,
        TranscriptionSegment,
    };

pub use crate::video::{generate_video, VideoFormat, VideoOptions, VideoProvider, VideoResult};

speech.rs

Read declaration text · 7 declaration entries

#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum AudioFormat {
    /// MP3 — widely supported lossy audio format.
    #[default]
    Mp3,
    /// WAV — uncompressed PCM audio.
    Wav,
    /// OGG — open container format, typically with Vorbis or Opus codec.
    Ogg,
    /// FLAC — lossless audio compression.
    Flac,
}

pub fn mime_type(&self) -> &'static str;

pub fn extension(&self) -> &'static str;

#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct SpeechOptions {
/// The voice identifier to use (e.g., "alloy", "echo", "nova").

/// `None` means the provider's default voice.

pub voice: Option<String>,
/// The playback speed multiplier (e.g., 0.5 for half speed, 2.0 for double speed).

/// `None` means the provider's default speed (typically 1.0).

pub speed: Option<f64>,
/// Output audio format.

pub format: AudioFormat
}

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct SpeechResult {
/// The raw audio bytes in the format specified by `mime_type`.

pub audio: Vec<u8>,
/// The MIME type of the audio data (e.g., "audio/mpeg").

pub mime_type: String,
/// The duration of the generated audio in seconds, if available.

/// `None` if the provider does not report duration.

pub duration_seconds: Option<f64>
}

#[async_trait]
pub trait SpeechProvider: Send + Sync {
    /// Returns the model identifier (e.g., "tts-1", "tts-1-hd").
    fn model_id(&self) -> &str;

    /// Returns the provider name (e.g., "openai", "elevenlabs").
    fn provider_name(&self) -> &str;

    /// Synthesizes speech from text.
    ///
    /// # Arguments
    ///
    /// * `text` - The text to convert to speech.
    /// * `options` - Configuration for voice, speed, and output format.
    ///
    /// # Returns
    ///
    /// A [`SpeechResult`] containing the generated audio bytes and metadata.
    ///
    /// # Errors
    ///
    /// * [`ForgeMediaError::SpeechFailed`] -- if the provider returns an error.
    /// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
    async fn speak(
        &self,
        text: &str,
        options: &SpeechOptions,
    ) -> Result<SpeechResult, ForgeMediaError>;
}

pub async fn speak(
    provider: &dyn SpeechProvider,
    text: &str,
    options: &SpeechOptions,
) -> Result<SpeechResult, ForgeMediaError>;

transcription.rs

Read declaration text · 6 declaration entries

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct TranscriptionSegment {
/// The start time of this segment in seconds from the beginning of the audio.

pub start: f64,
/// The end time of this segment in seconds from the beginning of the audio.

pub end: f64,
/// The transcribed text for this segment.

pub text: String
}

pub fn duration(&self) -> f64;

#[derive(Debug, Clone, Default, Serialize, Deserialize)]
pub struct TranscriptionOptions {
/// An ISO 639-1 language code hint (e.g., "en", "fr", "de").

/// `None` means the provider will auto-detect the language.

pub language: Option<String>,
/// An optional prompt to guide the transcription model. Useful for providing

/// context about domain-specific terminology or expected content.

/// `None` means no guidance prompt.

pub prompt: Option<String>
}

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct TranscriptionResult {
/// The full transcribed text.

pub text: String,
/// The detected language as an ISO 639-1 code, if available.

pub language: Option<String>,
/// The total duration of the audio in seconds, if available.

pub duration_seconds: Option<f64>,
/// Time-aligned transcription segments. May be empty if the provider

/// does not support segmented output.

pub segments: Vec<TranscriptionSegment>
}

#[async_trait]
pub trait TranscriptionProvider: Send + Sync {
    /// Returns the model identifier (e.g., "whisper-1", "chirp-v2").
    fn model_id(&self) -> &str;

    /// Returns the provider name (e.g., "openai", "google").
    fn provider_name(&self) -> &str;

    /// Transcribes audio data to text.
    ///
    /// # Arguments
    ///
    /// * `audio` - Raw audio bytes. The provider determines acceptable formats
    ///   (e.g., WAV, MP3, FLAC, OGG).
    /// * `options` - Configuration for language hint and guidance prompt.
    ///
    /// # Returns
    ///
    /// A [`TranscriptionResult`] containing the transcribed text and metadata.
    ///
    /// # Errors
    ///
    /// * [`ForgeMediaError::TranscriptionFailed`] -- if the provider returns an error.
    /// * [`ForgeMediaError::UnsupportedFormat`] -- if the audio format is not supported.
    async fn transcribe(
        &self,
        audio: &[u8],
        options: &TranscriptionOptions,
    ) -> Result<TranscriptionResult, ForgeMediaError>;
}

pub async fn transcribe(
    provider: &dyn TranscriptionProvider,
    audio: &[u8],
    options: &TranscriptionOptions,
) -> Result<TranscriptionResult, ForgeMediaError>;

video.rs

Read declaration text · 7 declaration entries

#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default, Serialize, Deserialize)]
#[serde(rename_all = "lowercase")]
pub enum VideoFormat {
    /// MP4 — widely supported video container format (typically H.264/AAC).
    #[default]
    Mp4,
    /// WebM — open video format (typically VP8/VP9/Opus).
    Webm,
}

pub fn mime_type(&self) -> &'static str;

pub fn extension(&self) -> &'static str;

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct VideoOptions {
/// Desired width in pixels.

pub width: u32,
/// Desired height in pixels.

pub height: u32,
/// Desired video duration in seconds.

pub duration_seconds: f64,
/// Output video format.

pub format: VideoFormat
}

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct VideoResult {
/// The raw video bytes in the format specified by `mime_type`.

pub data: Vec<u8>,
/// The MIME type of the video data (e.g., "video/mp4").

pub mime_type: String,
/// The duration of the generated video in seconds.

pub duration_seconds: f64,
/// The width of the generated video in pixels.

pub width: u32,
/// The height of the generated video in pixels.

pub height: u32
}

#[async_trait]
pub trait VideoProvider: Send + Sync {
    /// Returns the model identifier (e.g., "sora-1", "runway-gen2").
    fn model_id(&self) -> &str;

    /// Returns the provider name (e.g., "openai", "runway").
    fn provider_name(&self) -> &str;

    /// Generates a video from a text prompt.
    ///
    /// # Arguments
    ///
    /// * `prompt` - The text description of the video to generate.
    /// * `options` - Configuration for output dimensions, duration, and format.
    ///
    /// # Returns
    ///
    /// A [`VideoResult`] containing the generated video bytes and metadata.
    ///
    /// # Errors
    ///
    /// * [`ForgeMediaError::VideoFailed`] -- if the provider returns an error.
    /// * [`ForgeMediaError::UnsupportedFormat`] -- if the requested format is not supported.
    async fn generate_video(
        &self,
        prompt: &str,
        options: &VideoOptions,
    ) -> Result<VideoResult, ForgeMediaError>;
}

pub async fn generate_video(
    provider: &dyn VideoProvider,
    prompt: &str,
    options: &VideoOptions,
) -> Result<VideoResult, ForgeMediaError>;

Continue

On this page