forge-web
Native web substrate for the Forge SDK — fetch, parse, extract, crawl, and compact web content
Native web substrate for the Forge SDK — fetch, parse, extract, crawl, and compact web content
Package contract
| Field | Value |
|---|---|
| Language | rust |
| Source version | 0.2.0 |
| Manifest | forge-rs/crates/forge-web/Cargo.toml |
| Source files | 13 |
| Evidence | Source reference; registry publication and runtime conformance are separate checks |
Import boundary
use forge_web;Use a source checkout or your verified private registry. Manifest coordinates identify the package; they do not establish that a public registry release exists.
Crate boundary
The following entries are taken from src/lib.rs. Feature conditions in the exact source still apply.
pub mod compact;
pub mod config;
pub mod crawl;
pub mod error;
pub mod extract;
#[cfg(not(target_arch = "wasm32"))]
pub mod fetch;
pub mod inspect;
pub mod markdown;
pub mod parse;
pub mod search;
pub mod tools;
pub mod types;
pub mod prelude;
pub use crate::config::WebSubstrateConfig;
pub use crate::error::{WebError, WebResult};
#[cfg(not(target_arch = "wasm32"))]
pub use crate::fetch::web_fetch;
pub use crate::tools::register_web_tools;
pub use crate::types::{
CompactedSite, SiteMap, SiteNode, WebContent, WebExtractQuery, WebExtractResult,
WebFetchRequest, WebFetchResponse, WebSearchResponse, WebSearchResult,
};Source reference
Download package reference JSON. Each original source file and generated declaration artifact has its own SHA-256 digest. Function bodies and constant values are omitted from downloads. These are source declaration inventories, not compiler-resolved rustdoc, TypeDoc, DocC, or Dokka output. Private modules can contain public declarations that are not reachable through the package boundary; consult the entry point before importing.
compact.rs
Read declaration text · 1 declaration entries
pub async fn web_compact_site(
config: &WebSubstrateConfig,
root_url: &str,
max_depth: u32,
max_pages: u32,
) -> WebResult<CompactedSite>;config.rs
Read declaration text · 1 declaration entries
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebSubstrateConfig {
/// Maximum response body size in bytes.
///
/// Responses exceeding this limit will return
/// [`WebError::ContentTooLarge`](crate::error::WebError::ContentTooLarge).
///
/// Default: 10,485,760 (10 MB).
pub max_fetch_size_bytes: u64,
/// Maximum number of HTTP redirects to follow per request.
///
/// Exceeding this limit returns
/// [`WebError::RedirectLimitExceeded`](crate::error::WebError::RedirectLimitExceeded).
///
/// Default: 5.
pub max_redirects: u32,
/// Request timeout in milliseconds.
///
/// If the server does not respond within this duration, the request
/// returns [`WebError::Timeout`](crate::error::WebError::Timeout).
///
/// Default: 30,000 (30 seconds).
pub request_timeout_ms: u64,
/// IP address ranges to block for SSRF protection.
///
/// Each entry is a CIDR notation string (e.g., `10.0.0.0/8`). Requests
/// whose resolved IP falls within any of these ranges will return
/// [`WebError::SsrfBlocked`](crate::error::WebError::SsrfBlocked).
///
/// Default: private (RFC 1918), loopback, link-local, and IPv6 private ranges.
pub blocked_ip_ranges: Vec<String>,
/// The `User-Agent` header sent with all outgoing requests.
///
/// Default: `"forge-web/0.1 (+https://github.com/l1fe-labs/forge)"`.
pub user_agent: String,
/// Whether to respect `robots.txt` directives when crawling.
///
/// When `true`, the substrate will fetch and honor `robots.txt` rules
/// before crawling a site. Disabling this is only appropriate for
/// authorized internal crawling.
///
/// Default: `true`.
pub respect_robots_txt: bool
}crawl.rs
Read declaration text · 1 declaration entries
pub async fn web_crawl(
config: &WebSubstrateConfig,
root_url: &str,
max_depth: u32,
max_pages: u32,
) -> WebResult<Vec<WebFetchResponse>>;error.rs
Read declaration text · 2 declaration entries
#[derive(Debug, Error)]
pub enum WebError {
/// An HTTP fetch operation failed.
///
/// This covers network errors, DNS resolution failures, TLS handshake
/// failures, and other transport-level problems.
#[error("fetch failed for '{url}': {reason}")]
FetchFailed {
/// The URL that was being fetched.
url: String,
/// What went wrong during the fetch.
reason: String,
},
/// HTML or document parsing failed.
///
/// Returned when the web content cannot be parsed into the expected
/// structured format or does not satisfy the operation's parser contract.
#[error("parse failed for '{url}': {reason}")]
ParseFailed {
/// The URL whose content could not be parsed.
url: String,
/// What went wrong during parsing.
reason: String,
},
/// Content extraction failed.
///
/// Returned when a CSS selector, XPath, or JSONPath query cannot be
/// executed against the parsed content.
#[error("extraction failed for query '{query}' on '{url}': {reason}")]
ExtractionFailed {
/// The URL whose content was being queried.
url: String,
/// The extraction query that failed.
query: String,
/// What went wrong during extraction.
reason: String,
},
/// A web search operation failed.
///
/// Returned when a search query cannot be validated, a search endpoint
/// cannot be queried, or the response cannot be parsed into results.
#[error("search failed for query '{query}': {reason}")]
SearchFailed {
/// The search query.
query: String,
/// What went wrong during the search.
reason: String,
},
/// A fetch was blocked because the resolved IP address falls within a
/// private, loopback, or link-local range (SSRF protection).
///
/// This is a security control. The blocked IP ranges are configured via
/// [`WebSubstrateConfig::blocked_ip_ranges`](crate::config::WebSubstrateConfig).
#[error(
"SSRF blocked: '{url}' resolved to blocked IP {ip} (private/loopback/link-local range)"
)]
SsrfBlocked {
/// The URL that was being fetched.
url: String,
/// The IP address that triggered the block.
ip: String,
},
/// The maximum number of HTTP redirects was exceeded.
///
/// Configure the limit via
/// [`WebSubstrateConfig::max_redirects`](crate::config::WebSubstrateConfig).
#[error("redirect limit exceeded for '{url}': followed {count} redirects (max {max})")]
RedirectLimitExceeded {
/// The original URL that was being fetched.
url: String,
/// The number of redirects followed before the limit was hit.
count: u32,
/// The configured maximum number of redirects.
max: u32,
},
/// The response body exceeds the configured maximum size.
///
/// Configure the limit via
/// [`WebSubstrateConfig::max_fetch_size_bytes`](crate::config::WebSubstrateConfig).
#[error(
"content too large for '{url}': response size {size} bytes exceeds limit of {max} bytes"
)]
ContentTooLarge {
/// The URL whose response was too large.
url: String,
/// The actual (or estimated) response size in bytes.
size: u64,
/// The configured maximum size in bytes.
max: u64,
},
/// The provided URL is invalid or cannot be parsed.
///
/// Check that the URL includes a scheme (`http://` or `https://`),
/// a valid host, and well-formed path components.
#[error("invalid URL '{url}': {reason}")]
InvalidUrl {
/// The URL string that failed validation.
url: String,
/// What is wrong with the URL.
reason: String,
},
/// The HTTP request timed out.
///
/// Configure the timeout via
/// [`WebSubstrateConfig::request_timeout_ms`](crate::config::WebSubstrateConfig).
#[error("request timed out for '{url}' after {timeout_ms}ms")]
Timeout {
/// The URL that timed out.
url: String,
/// The configured timeout in milliseconds.
timeout_ms: u64,
},
/// A boundary contract denied the operation.
///
/// This occurs when the web substrate detects a security policy violation
/// such as a cross-scheme redirect downgrade (HTTPS to HTTP).
#[error("boundary contract denied for '{url}': {reason}")]
BoundaryContractDenied {
/// The URL involved in the denied operation.
url: String,
/// Why the boundary contract was violated.
reason: String,
},
}
pub type WebResult<T> = Result<T, WebError>;extract.rs
Read declaration text · 1 declaration entries
pub fn web_extract(
content: &str,
query: &WebExtractQuery,
url: &str,
) -> WebResult<WebExtractResult>;fetch.rs
Read declaration text · 1 declaration entries
pub async fn web_fetch(
config: &WebSubstrateConfig,
request: &WebFetchRequest,
) -> WebResult<WebFetchResponse>;inspect.rs
Read declaration text · 1 declaration entries
pub async fn web_inspect_site(
config: &WebSubstrateConfig,
root_url: &str,
max_depth: u32,
max_pages: u32,
) -> WebResult<SiteMap>;lib.rs
Read declaration text · 18 declaration entries
pub mod compact;
pub mod config;
pub mod crawl;
pub mod error;
pub mod extract;
#[cfg(not(target_arch = "wasm32"))]
pub mod fetch;
pub mod inspect;
pub mod markdown;
pub mod parse;
pub mod search;
pub mod tools;
pub mod types;
pub mod prelude;
pub use crate::config::WebSubstrateConfig;
pub use crate::error::{WebError, WebResult};
#[cfg(not(target_arch = "wasm32"))]
pub use crate::fetch::web_fetch;
pub use crate::tools::register_web_tools;
pub use crate::types::{
CompactedSite, SiteMap, SiteNode, WebContent, WebExtractQuery, WebExtractResult,
WebFetchRequest, WebFetchResponse, WebSearchResponse, WebSearchResult,
};markdown.rs
Read declaration text · 1 declaration entries
pub fn web_to_markdown(html: &str, url: &str) -> WebResult<WebContent>;parse.rs
Read declaration text · 1 declaration entries
pub fn web_parse(html: &str, url: &str) -> WebResult<WebContent>;search.rs
Read declaration text · 2 declaration entries
pub const DEFAULT_SEARCH_ENDPOINT: &str;
pub async fn web_search(
config: &WebSubstrateConfig,
query: &str,
num_results: u32,
) -> WebResult<WebSearchResponse>;tools.rs
Read declaration text · 3 declaration entries
pub const WEB_TOOL_NAMES: &[&str];
pub fn web_tool_definitions() -> Vec<ToolDefinition>;
pub fn register_web_tools(registry: &mut ToolRegistry) -> Result<(), ForgeToolError>;types.rs
Read declaration text · 18 declaration entries
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub enum HttpMethod {
/// HTTP GET request.
#[serde(rename = "GET")]
Get,
/// HTTP POST request.
#[serde(rename = "POST")]
Post,
/// HTTP HEAD request.
#[serde(rename = "HEAD")]
Head,
}
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebFetchRequest {
/// The target URL to fetch.
pub url: String,
/// The HTTP method to use.
pub method: HttpMethod,
/// HTTP headers to include in the request.
///
/// Keys are header names, values are header values. Uses `BTreeMap` for
/// deterministic serialization.
pub headers: BTreeMap<String, String>,
/// Request timeout in milliseconds. Overrides the config default if set.
pub timeout_ms: Option<u64>,
/// Whether to follow HTTP redirects.
pub follow_redirects: bool,
/// Maximum number of redirects to follow. Overrides the config default if set.
pub max_redirects: Option<u32>,
/// Optional request body (for POST requests).
#[serde(skip_serializing_if = "Option::is_none")]
pub body: Option<String>
}
pub fn get(url: &str) -> Result<Self, WebError>;
pub fn post(url: &str) -> Result<Self, WebError>;
pub fn head(url: &str) -> Result<Self, WebError>;
pub fn with_header(mut self, name: impl Into<String>, value: impl Into<String>) -> Self;
pub fn with_body(mut self, body: impl Into<String>) -> Self;
pub fn with_timeout(mut self, timeout_ms: u64) -> Self;
#[derive(Debug, Clone, Serialize, Deserialize)]
pub struct WebFetchResponse {
/// The HTTP status code (e.g., 200, 404, 500).
pub status: u16,
/// Response headers. Uses `BTreeMap` for deterministic serialization.
pub headers: BTreeMap<String, String>,
/// The response body as a string.
///
/// Binary responses are base64-encoded. Non-UTF-8 text responses use
/// lossy conversion.
pub body: String,
/// The detected content type from the `Content-Type` header.
///
/// `None` if no `Content-Type` header is present.
pub content_type: Option<String>,
/// The final URL after following any redirects.
///
/// Matches the request URL if no redirects occurred.
pub final_url: String
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(tag = "type", content = "data")]
pub enum WebContent {
/// Raw HTML content.
Html(String),
/// Markdown-converted content.
Markdown(String),
/// Plain text content (HTML tags stripped).
PlainText(String),
/// Structured JSON content.
Json(serde_json::Value),
/// Binary content (e.g., images, PDFs).
Binary(Vec<u8>),
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
#[serde(tag = "type", content = "expression")]
pub enum WebExtractQuery {
/// A CSS selector query (e.g., `div.content > p`).
CssSelector(String),
/// An XPath query (e.g., `//div[@class='content']/p`).
XPath(String),
/// A JSONPath query (e.g., `$.data.items[*].name`).
JsonPath(String),
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct WebExtractResult {
/// The matched elements or values as strings.
pub matches: Vec<String>
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct SiteNode {
/// The URL of this page.
pub url: String,
/// The page title extracted from the `<title>` tag, if available.
pub title: Option<String>,
/// URLs linked from this page.
pub links: Vec<String>,
/// The crawl depth at which this node was discovered (0 = root).
pub depth: u32
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct SiteMap {
/// The root URL that the crawl started from.
pub root: String,
/// All discovered site nodes, in breadth-first order.
pub nodes: Vec<SiteNode>
}
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
pub struct CompactedSite {
/// Map of URL to compacted Markdown content. Uses `BTreeMap` for
/// deterministic serialization order.
pub pages: BTreeMap<String, String>,
/// Estimated total token count across all pages.
///
/// Uses a rough heuristic of ~4 characters per token.
pub total_tokens_estimate: u64
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct WebSearchResult {
/// Human-readable search result title.
pub title: String,
/// Canonical result URL.
pub url: String,
/// Optional search-provider snippet or summary.
pub snippet: Option<String>
}
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
pub struct WebSearchResponse {
/// Original user query after trimming leading/trailing whitespace.
pub query: String,
/// Search results in provider rank order.
pub results: Vec<WebSearchResult>
}
pub fn validate_url(url: &str) -> Result<String, WebError>;