Local AI 3D Model Generation Pipeline
Independent technical experiment
- My role
- AI / ML Experimentation
- Team
- Solo project
- Tools
- Trellis2
- GGUF
- PyTorch
- Transformers
- Hugging Face
- Modly
- Timeline
May 2026 — Experimental Prototype
- Description
An experimental local AI pipeline exploring image-to-3D model generation using modern open-source vision and generative-model tooling.
- Context
I explored whether modern image-to-3D generation workflows could be run locally rather than depending entirely on hosted AI services. The experiment was primarily technical: understanding the model dependencies, inference environment, local hardware requirements, model formats, and practical barriers involved in assembling an open-source generative pipeline.
Click around...
Challenge
Modern generative-AI projects often appear simple when viewed through a hosted demo but are considerably more involved when reproduced locally.
The pipeline depended on several model components, Python libraries, model weights, compatible environments, and external repositories.
A single inaccessible dependency could prevent the entire inference chain from running.
Constraints
The experiment encountered a gated Hugging Face dependency, facebook/dinov3-vitl16-pretrain-lvd1689m, which required authenticated access.
That prevented the attempted configuration from completing successfully.
For that reason, this project should be presented as an experiment in local AI infrastructure rather than a completed production system.
Research
I investigated Trellis2 and the surrounding ecosystem for local image-to-3D generation.
The work involved understanding GGUF model formats, PyTorch and Transformers dependencies, Hugging Face model distribution, and the model-loading requirements of the broader pipeline.
Iterations
I assembled the local environment and worked through the dependencies required to launch the pipeline.
The experiment progressed to model-loading and execution attempts but was blocked by the gated DINOv3 dependency.
Rather than presenting the experiment as a finished application, I retained it as a useful investigation into the practical complexity of running multimodal generative models locally.
Key Features
- Local generative-AI experimentation
- Image-to-3D pipeline investigation
- Trellis2
- GGUF model tooling
- PyTorch
- Transformers
- Hugging Face model loading
- Dependency and environment configuration
Final Deliverable
A partially implemented local image-to-3D inference pipeline and technical investigation into the infrastructure required to operate modern generative models outside hosted platforms.
Takeaway
The project demonstrated the gap between using an AI model and operating one.
Local inference involves model licensing and access, hardware compatibility, dependency management, memory requirements, model formats, and infrastructure concerns that are largely hidden by hosted APIs.
Even though the final generation pipeline was blocked by a gated dependency, the experiment provided useful hands-on experience with that lower layer of the AI stack.