Jonas Ciplickas — Scale Venture Partners - VLMs are finally giving AEC its 'AI will change everything' moment - February 2026 logo

Jonas Ciplickas — Scale Venture Partners - VLMs are finally giving AEC its 'AI will change everything' moment - February 2026

Free

VLMs are finally giving AEC its 'AI will change everything' moment

FreeFree tier
Inputs: text, imageOutputs: text
Type
Open Source

About Jonas Ciplickas — Scale Venture Partners - VLMs are finally giving AEC its 'AI will change everything' moment - February 2026

This article by Jonas Ciplickas of Scale Venture Partners, published in February 2026, argues that Vision-Language Models (VLMs) are finally enabling transformative AI applications in the Architecture, Engineering, and Construction (AEC) industry. It explains how VLMs extend LLM reasoning to images, allowing for semantic, spatial, geometric, and numeric reasoning across multimodal inputs. The piece outlines how VLMs overcome limitations of prior cloud software and LLMs in AEC, enabling end-to-end workflows in areas like documentation creation, permitting, takeoffs, estimation, procurement, and project management. It also discusses the technical workings of VLMs, their rapid advancement (inheriting LLM improvements every ~6 months), and their inherent weaknesses, positioning them as the catalyst for a new wave of $10B+ AEC software companies.

Key Features

Accepts both text and image inputs for reasoning and execution
Combines semantic, spatial, geometric, and numeric reasoning in single context window
Inherits LLM capabilities: longer task horizons, tool calling, orchestration
Enables end-to-end workflows by pairing VLM with deterministic tools
Provides interpretation and reasoning beyond simple classification

Pros & Cons

Pros
  • Enables AI to understand both text and image-based inputs critical for physical-world industries
  • Rapidly improves every ~6 months as foundation models advance
  • Unlocks entirely new product categories and massive market expansion for AEC software
Cons
  • Inherits LLM weaknesses such as limited precision and reliability in high-stakes tasks
  • Requires pairing with deterministic tools to compensate for shortcomings
  • Early applications still face integration challenges with existing AEC workflows

Best For

Documentation creation in constructionPermitting workflowsTakeoffs, estimation, and procurementProject management across construction lifecycleGeneral AEC software applications requiring multimodal inputs

FAQ

What is a Vision-Language Model (VLM)?
A VLM takes in an image plus a text prompt and outputs a response, extending LLM reasoning to images. It breaks images into chunks, translates them into tokens, and combines image and text tokens into a single context window for the model to respond.
How do VLMs differ from LLMs in AEC applications?
LLMs struggled in AEC because construction requires understanding complex, heterogeneous, image-based data beyond text. VLMs enable reasoning over both text and images, allowing workflows like takeoffs, permitting, and project management to become end-to-end AI-driven.