Forge documentation
Library referenceRust

forge-web

Native web substrate for the Forge SDK — fetch, parse, extract, crawl, and compact web content

Native web substrate for the Forge SDK — fetch, parse, extract, crawl, and compact web content

Package contract

FieldValue
Languagerust
Source version0.2.0
Manifestforge-rs/crates/forge-web/Cargo.toml
Source files13
EvidenceSource reference; registry publication and runtime conformance are separate checks

Import boundary

use forge_web;

Use a source checkout or your verified private registry. Manifest coordinates identify the package; they do not establish that a public registry release exists.

Crate boundary

The following entries are taken from src/lib.rs. Feature conditions in the exact source still apply.

pub mod compact;

pub mod config;

pub mod crawl;

pub mod error;

pub mod extract;

#[cfg(not(target_arch = "wasm32"))]
pub mod fetch;

pub mod inspect;

pub mod markdown;

pub mod parse;

pub mod search;

pub mod tools;

pub mod types;

pub mod prelude;

pub use crate::config::WebSubstrateConfig;

pub use crate::error::{WebError, WebResult};

#[cfg(not(target_arch = "wasm32"))]
pub use crate::fetch::web_fetch;

pub use crate::tools::register_web_tools;

pub use crate::types::{
        CompactedSite, SiteMap, SiteNode, WebContent, WebExtractQuery, WebExtractResult,
        WebFetchRequest, WebFetchResponse, WebSearchResponse, WebSearchResult,
    };

Source reference

Download package reference JSON. Each original source file and generated declaration artifact has its own SHA-256 digest. Function bodies and constant values are omitted from downloads. These are source declaration inventories, not compiler-resolved rustdoc, TypeDoc, DocC, or Dokka output. Private modules can contain public declarations that are not reachable through the package boundary; consult the entry point before importing.

compact.rs

Read declaration text · 1 declaration entries

pub async fn web_compact_site(
    config: &WebSubstrateConfig,
    root_url: &str,
    max_depth: u32,
    max_pages: u32,
) -> WebResult<CompactedSite>;

config.rs

Read declaration text · 1 declaration entries

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebSubstrateConfig {
/// Maximum response body size in bytes.

///

/// Responses exceeding this limit will return

/// [`WebError::ContentTooLarge`](crate::error::WebError::ContentTooLarge).

///

/// Default: 10,485,760 (10 MB).

pub max_fetch_size_bytes: u64,
/// Maximum number of HTTP redirects to follow per request.

///

/// Exceeding this limit returns

/// [`WebError::RedirectLimitExceeded`](crate::error::WebError::RedirectLimitExceeded).

///

/// Default: 5.

pub max_redirects: u32,
/// Request timeout in milliseconds.

///

/// If the server does not respond within this duration, the request

/// returns [`WebError::Timeout`](crate::error::WebError::Timeout).

///

/// Default: 30,000 (30 seconds).

pub request_timeout_ms: u64,
/// IP address ranges to block for SSRF protection.

///

/// Each entry is a CIDR notation string (e.g., `10.0.0.0/8`). Requests

/// whose resolved IP falls within any of these ranges will return

/// [`WebError::SsrfBlocked`](crate::error::WebError::SsrfBlocked).

///

/// Default: private (RFC 1918), loopback, link-local, and IPv6 private ranges.

pub blocked_ip_ranges: Vec<String>,
/// The `User-Agent` header sent with all outgoing requests.

///

/// Default: `"forge-web/0.1 (+https://github.com/l1fe-labs/forge)"`.

pub user_agent: String,
/// Whether to respect `robots.txt` directives when crawling.

///

/// When `true`, the substrate will fetch and honor `robots.txt` rules

/// before crawling a site. Disabling this is only appropriate for

/// authorized internal crawling.

///

/// Default: `true`.

pub respect_robots_txt: bool
}

crawl.rs

Read declaration text · 1 declaration entries

pub async fn web_crawl(
    config: &WebSubstrateConfig,
    root_url: &str,
    max_depth: u32,
    max_pages: u32,
) -> WebResult<Vec<WebFetchResponse>>;

error.rs

Read declaration text · 2 declaration entries

#[derive(Debug, Error)]
pub enum WebError {
    /// An HTTP fetch operation failed.
    ///
    /// This covers network errors, DNS resolution failures, TLS handshake
    /// failures, and other transport-level problems.
    #[error("fetch failed for '{url}': {reason}")]
    FetchFailed {
        /// The URL that was being fetched.
        url: String,
        /// What went wrong during the fetch.
        reason: String,
    },

    /// HTML or document parsing failed.
    ///
    /// Returned when the web content cannot be parsed into the expected
    /// structured format or does not satisfy the operation's parser contract.
    #[error("parse failed for '{url}': {reason}")]
    ParseFailed {
        /// The URL whose content could not be parsed.
        url: String,
        /// What went wrong during parsing.
        reason: String,
    },

    /// Content extraction failed.
    ///
    /// Returned when a CSS selector, XPath, or JSONPath query cannot be
    /// executed against the parsed content.
    #[error("extraction failed for query '{query}' on '{url}': {reason}")]
    ExtractionFailed {
        /// The URL whose content was being queried.
        url: String,
        /// The extraction query that failed.
        query: String,
        /// What went wrong during extraction.
        reason: String,
    },

    /// A web search operation failed.
    ///
    /// Returned when a search query cannot be validated, a search endpoint
    /// cannot be queried, or the response cannot be parsed into results.
    #[error("search failed for query '{query}': {reason}")]
    SearchFailed {
        /// The search query.
        query: String,
        /// What went wrong during the search.
        reason: String,
    },

    /// A fetch was blocked because the resolved IP address falls within a
    /// private, loopback, or link-local range (SSRF protection).
    ///
    /// This is a security control. The blocked IP ranges are configured via
    /// [`WebSubstrateConfig::blocked_ip_ranges`](crate::config::WebSubstrateConfig).
    #[error(
        "SSRF blocked: '{url}' resolved to blocked IP {ip} (private/loopback/link-local range)"
    )]
    SsrfBlocked {
        /// The URL that was being fetched.
        url: String,
        /// The IP address that triggered the block.
        ip: String,
    },

    /// The maximum number of HTTP redirects was exceeded.
    ///
    /// Configure the limit via
    /// [`WebSubstrateConfig::max_redirects`](crate::config::WebSubstrateConfig).
    #[error("redirect limit exceeded for '{url}': followed {count} redirects (max {max})")]
    RedirectLimitExceeded {
        /// The original URL that was being fetched.
        url: String,
        /// The number of redirects followed before the limit was hit.
        count: u32,
        /// The configured maximum number of redirects.
        max: u32,
    },

    /// The response body exceeds the configured maximum size.
    ///
    /// Configure the limit via
    /// [`WebSubstrateConfig::max_fetch_size_bytes`](crate::config::WebSubstrateConfig).
    #[error(
        "content too large for '{url}': response size {size} bytes exceeds limit of {max} bytes"
    )]
    ContentTooLarge {
        /// The URL whose response was too large.
        url: String,
        /// The actual (or estimated) response size in bytes.
        size: u64,
        /// The configured maximum size in bytes.
        max: u64,
    },

    /// The provided URL is invalid or cannot be parsed.
    ///
    /// Check that the URL includes a scheme (`http://` or `https://`),
    /// a valid host, and well-formed path components.
    #[error("invalid URL '{url}': {reason}")]
    InvalidUrl {
        /// The URL string that failed validation.
        url: String,
        /// What is wrong with the URL.
        reason: String,
    },

    /// The HTTP request timed out.
    ///
    /// Configure the timeout via
    /// [`WebSubstrateConfig::request_timeout_ms`](crate::config::WebSubstrateConfig).
    #[error("request timed out for '{url}' after {timeout_ms}ms")]
    Timeout {
        /// The URL that timed out.
        url: String,
        /// The configured timeout in milliseconds.
        timeout_ms: u64,
    },

    /// A boundary contract denied the operation.
    ///
    /// This occurs when the web substrate detects a security policy violation
    /// such as a cross-scheme redirect downgrade (HTTPS to HTTP).
    #[error("boundary contract denied for '{url}': {reason}")]
    BoundaryContractDenied {
        /// The URL involved in the denied operation.
        url: String,
        /// Why the boundary contract was violated.
        reason: String,
    },
}

pub type WebResult<T> = Result<T, WebError>;

extract.rs

Read declaration text · 1 declaration entries

pub fn web_extract(
    content: &str,
    query: &WebExtractQuery,
    url: &str,
) -> WebResult<WebExtractResult>;

fetch.rs

Read declaration text · 1 declaration entries

pub async fn web_fetch(
    config: &WebSubstrateConfig,
    request: &WebFetchRequest,
) -> WebResult<WebFetchResponse>;

inspect.rs

Read declaration text · 1 declaration entries

pub async fn web_inspect_site(
    config: &WebSubstrateConfig,
    root_url: &str,
    max_depth: u32,
    max_pages: u32,
) -> WebResult<SiteMap>;

lib.rs

Read declaration text · 18 declaration entries

pub mod compact;

pub mod config;

pub mod crawl;

pub mod error;

pub mod extract;

#[cfg(not(target_arch = "wasm32"))]
pub mod fetch;

pub mod inspect;

pub mod markdown;

pub mod parse;

pub mod search;

pub mod tools;

pub mod types;

pub mod prelude;

pub use crate::config::WebSubstrateConfig;

pub use crate::error::{WebError, WebResult};

#[cfg(not(target_arch = "wasm32"))]
pub use crate::fetch::web_fetch;

pub use crate::tools::register_web_tools;

pub use crate::types::{
        CompactedSite, SiteMap, SiteNode, WebContent, WebExtractQuery, WebExtractResult,
        WebFetchRequest, WebFetchResponse, WebSearchResponse, WebSearchResult,
    };

markdown.rs

Read declaration text · 1 declaration entries

pub fn web_to_markdown(html: &str, url: &str) -> WebResult<WebContent>;

parse.rs

Read declaration text · 1 declaration entries

pub fn web_parse(html: &str, url: &str) -> WebResult<WebContent>;

search.rs

Read declaration text · 2 declaration entries

pub const DEFAULT_SEARCH_ENDPOINT: &str;

pub async fn web_search(
    config: &WebSubstrateConfig,
    query: &str,
    num_results: u32,
) -> WebResult<WebSearchResponse>;

tools.rs

Read declaration text · 3 declaration entries

pub const WEB_TOOL_NAMES: &[&str];

pub fn web_tool_definitions() -> Vec<ToolDefinition>;

pub fn register_web_tools(registry: &mut ToolRegistry) -> Result<(), ForgeToolError>;

types.rs

Read declaration text · 18 declaration entries

#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub enum HttpMethod {
    /// HTTP GET request.
    #[serde(rename = "GET")]
    Get,
    /// HTTP POST request.
    #[serde(rename = "POST")]
    Post,
    /// HTTP HEAD request.
    #[serde(rename = "HEAD")]
    Head,
}

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebFetchRequest {
/// The target URL to fetch.

pub url: String,
/// The HTTP method to use.

pub method: HttpMethod,
/// HTTP headers to include in the request.

///

/// Keys are header names, values are header values. Uses `BTreeMap` for

/// deterministic serialization.

pub headers: BTreeMap<String, String>,
/// Request timeout in milliseconds. Overrides the config default if set.

pub timeout_ms: Option<u64>,
/// Whether to follow HTTP redirects.

pub follow_redirects: bool,
/// Maximum number of redirects to follow. Overrides the config default if set.

pub max_redirects: Option<u32>,
/// Optional request body (for POST requests).

#[serde(skip_serializing_if = "Option::is_none")]
pub body: Option<String>
}

pub fn get(url: &str) -> Result<Self, WebError>;

pub fn post(url: &str) -> Result<Self, WebError>;

pub fn head(url: &str) -> Result<Self, WebError>;

pub fn with_header(mut self, name: impl Into<String>, value: impl Into<String>) -> Self;

pub fn with_body(mut self, body: impl Into<String>) -> Self;

pub fn with_timeout(mut self, timeout_ms: u64) -> Self;

#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebFetchResponse {
/// The HTTP status code (e.g., 200, 404, 500).

pub status: u16,
/// Response headers. Uses `BTreeMap` for deterministic serialization.

pub headers: BTreeMap<String, String>,
/// The response body as a string.

///

/// Binary responses are base64-encoded. Non-UTF-8 text responses use

/// lossy conversion.

pub body: String,
/// The detected content type from the `Content-Type` header.

///

/// `None` if no `Content-Type` header is present.

pub content_type: Option<String>,
/// The final URL after following any redirects.

///

/// Matches the request URL if no redirects occurred.

pub final_url: String
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(tag = "type", content = "data")]
pub enum WebContent {
    /// Raw HTML content.
    Html(String),
    /// Markdown-converted content.
    Markdown(String),
    /// Plain text content (HTML tags stripped).
    PlainText(String),
    /// Structured JSON content.
    Json(serde_json::Value),
    /// Binary content (e.g., images, PDFs).
    Binary(Vec<u8>),
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(tag = "type", content = "expression")]
pub enum WebExtractQuery {
    /// A CSS selector query (e.g., `div.content > p`).
    CssSelector(String),
    /// An XPath query (e.g., `//div[@class='content']/p`).
    XPath(String),
    /// A JSONPath query (e.g., `$.data.items[*].name`).
    JsonPath(String),
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct WebExtractResult {
/// The matched elements or values as strings.

pub matches: Vec<String>
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct SiteNode {
/// The URL of this page.

pub url: String,
/// The page title extracted from the `<title>` tag, if available.

pub title: Option<String>,
/// URLs linked from this page.

pub links: Vec<String>,
/// The crawl depth at which this node was discovered (0 = root).

pub depth: u32
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct SiteMap {
/// The root URL that the crawl started from.

pub root: String,
/// All discovered site nodes, in breadth-first order.

pub nodes: Vec<SiteNode>
}

#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct CompactedSite {
/// Map of URL to compacted Markdown content. Uses `BTreeMap` for

/// deterministic serialization order.

pub pages: BTreeMap<String, String>,
/// Estimated total token count across all pages.

///

/// Uses a rough heuristic of ~4 characters per token.

pub total_tokens_estimate: u64
}

#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct WebSearchResult {
/// Human-readable search result title.

pub title: String,
/// Canonical result URL.

pub url: String,
/// Optional search-provider snippet or summary.

pub snippet: Option<String>
}

#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct WebSearchResponse {
/// Original user query after trimming leading/trailing whitespace.

pub query: String,
/// Search results in provider rank order.

pub results: Vec<WebSearchResult>
}

pub fn validate_url(url: &str) -> Result<String, WebError>;

Continue

On this page