Inference.net: Schematron V2 Turbo
by Inference Nettext->text1 endpoint
Overview
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather than through system or user prompts.
Capabilities
Text generation
Structured output
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed
Modalities
Input
Text
Output
Text
Technical Specifications
Context Window
128K tokens
Max Output
8.2K tokens
Tokenizer
Other
Uptime (24h)
10000.0%
Supported Parameters
frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyresponse_formatseedstopstructured_outputstemperaturetop_ktop_p
Pricing
Per 1M tokens
Input$0.03
Output$0.15
Cache read$0.03
Added
September 12, 2026
Last synced 9/21/2026