Inference.net: Schematron V2 Turbo

by Inference Nettext->text1 endpoint

Overview

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather than through system or user prompts.

Capabilities

Text generation
Structured output
Prompt caching
Long context
Top-P sampling
Stop sequences
Deterministic seed

Modalities

Input
Text
Output
Text

Technical Specifications

Context Window
128K tokens
Max Output
8.2K tokens
Tokenizer
Other
Uptime (24h)
10000.0%

Supported Parameters

frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyresponse_formatseedstopstructured_outputstemperaturetop_ktop_p

Pricing

Per 1M tokens
Input$0.03
Output$0.15
Cache read$0.03
Added
September 12, 2026
Last synced 9/21/2026