Local AI 3D Model Generation Pipeline

Independent technical experiment


My role
  • AI / ML Experimentation
Team
  • Solo project
Tools
  • Trellis2
  • GGUF
  • PyTorch
  • Transformers
  • Hugging Face
  • Modly
Timeline

May 2026 — Experimental Prototype

Description

An experimental local AI pipeline exploring image-to-3D model generation using modern open-source vision and generative-model tooling.

Context

I explored whether modern image-to-3D generation workflows could be run locally rather than depending entirely on hosted AI services. The experiment was primarily technical: understanding the model dependencies, inference environment, local hardware requirements, model formats, and practical barriers involved in assembling an open-source generative pipeline.

Click around...

Challenge

Modern generative-AI projects often appear simple when viewed through a hosted demo but are considerably more involved when reproduced locally.

The pipeline depended on several model components, Python libraries, model weights, compatible environments, and external repositories.

A single inaccessible dependency could prevent the entire inference chain from running.

Constraints

The experiment encountered a gated Hugging Face dependency, facebook/dinov3-vitl16-pretrain-lvd1689m, which required authenticated access.

That prevented the attempted configuration from completing successfully.

For that reason, this project should be presented as an experiment in local AI infrastructure rather than a completed production system.

Research

I investigated Trellis2 and the surrounding ecosystem for local image-to-3D generation.

The work involved understanding GGUF model formats, PyTorch and Transformers dependencies, Hugging Face model distribution, and the model-loading requirements of the broader pipeline.

Iterations

I assembled the local environment and worked through the dependencies required to launch the pipeline.

The experiment progressed to model-loading and execution attempts but was blocked by the gated DINOv3 dependency.

Rather than presenting the experiment as a finished application, I retained it as a useful investigation into the practical complexity of running multimodal generative models locally.

Key Features

  • Local generative-AI experimentation
  • Image-to-3D pipeline investigation
  • Trellis2
  • GGUF model tooling
  • PyTorch
  • Transformers
  • Hugging Face model loading
  • Dependency and environment configuration

Final Deliverable

A partially implemented local image-to-3D inference pipeline and technical investigation into the infrastructure required to operate modern generative models outside hosted platforms.

Takeaway

The project demonstrated the gap between using an AI model and operating one.

Local inference involves model licensing and access, hardware compatibility, dependency management, memory requirements, model formats, and infrastructure concerns that are largely hidden by hosted APIs.

Even though the final generation pipeline was blocked by a gated dependency, the experiment provided useful hands-on experience with that lower layer of the AI stack.