Most AI developers were focused on one thing Getting an LLM to work. But, that is no longer enough.
The real engineering challenge is increasingly about:
Serving models efficiently
Running multiple AI agents at the same time
Building self-hosted AI interfaces
Creating reliable AI applications
Connecting models to production workflows
Supporting hundreds of models without rewriting your stack
In other words, We are moving from AI experimentation to AI infrastructure.
For Week 19 of my Open Source GitHub Repository Series, I explored five repositories that represent very different layers of that infrastructure:
vLLM β High-throughput LLM inference and serving
cmux β A developer terminal designed around parallel AI coding agents
Open WebUI β A self-hosted AI platform for models, agents, RAG, tools, and workflows
Dify β An open-source platform for building agentic workflows and AI applications
Hugging Face Transformers β The model-definition framework powering a huge portion of the modern AI ecosystem
What makes this weekβs list especially interesting is that these projects operate at different levels of the stack.
Some help you run the model.
Some help you work with the model.
Some help you build applications around the model.
And one helps connect the broader model ecosystem together.
Letβs dive in.
1. vLLM: The Inference Engine Behind Production-Scale AI
Repository: vllm-project/vllm
There is a big difference between running a model and serving a model efficiently.
Downloading an LLM and getting one response is relatively easy. The challenge begins when thousands of users start sending requests.
Now you care about:
throughput
latency
memory usage
batching
GPU utilization
concurrency
distributed inference
model compatibility
That is where vLLM becomes important.
What isΒ vLLM?
vLLM is an open-source library for fast and easy-to-use LLM inference and serving.
The project focuses heavily on throughput and memory efficiency and includes techniques such as PagedAttention, continuous batching, chunked prefill, prefix caching, optimized attention kernels, quantization, speculative decoding, and distributed parallelism.
The project also provides an OpenAI-compatible API server, along with Anthropic Messages API and gRPC support.
That makes vLLM much more than a research library.
It can become the serving layer sitting between your application and the model.
Think of the architecture as:
Your Application
β
vLLM
β
LLM Model
β
GPU / TPUInstead of every application figuring out how to efficiently execute models, vLLM handles the serving layer.
Why Developers LoveΒ It
One of the hardest parts of LLM infrastructure is making expensive hardware useful. A GPU that spends most of its time waiting for requests is wasted infrastructure.
vLLM focuses directly on increasing serving efficiency.
Its README highlights state-of-the-art serving throughput, efficient KV-cache memory management, continuous batching, prefix caching, quantization support, optimized kernels, speculative decoding, and distributed inference capabilities.
It also supports a broad range of model architectures and hardware.
The repository currently describes support for 200+ model architectures on Hugging Face, including decoder-only LLMs, Mixture-of-Experts models, hybrid architectures, multimodal models, embedding/retrieval models, and reward/classification models. It also lists support across NVIDIA, AMD, Intel and other hardware ecosystems.
That makes it especially interesting for teams building their own AI infrastructure.
Real Development UseΒ Cases
Self-Hosted LLM APIs: Run an open-source model behind an API that your applications can consume.
AI SaaS Platforms: Serve models for multiple customers while managing concurrency and infrastructure efficiently.
Internal AI Platforms: Provide a common model-serving layer for multiple internal applications.
RAG Systems: Deploy the generation model behind a retrieval pipeline.
Agent Platforms: Serve the models used by coding agents and autonomous workflows.
Model Experimentation: Quickly switch between supported models without rebuilding the whole serving layer.
Productivity Impact
Developers donβt want to spend their time writing low-level inference infrastructure.
They want to build products.
vLLM provides a mature serving layer so teams can focus more on:
application logic
agent behavior
evaluation
product features
and less on reinventing model-serving infrastructure.
The important distinction is this: Hugging Face Transformers helps define and use models. vLLM helps serve those models efficiently in real systems.
That difference becomes extremely important once your AI application leaves the prototype stage.
2. cmux: A Terminal Built for Developers Running AI Agents inΒ Parallel
Repository: manaflow-ai/cmux
Modern coding workflows are becoming strangely crowded.
You might have:
Claude Code
Codex
OpenCode
a local terminal
a development server
browser tools
Git
multiple repositories
several tasks running simultaneously
Open ten terminal tabs and eventually the terminal becomes the problem.
Which agent is waiting?
Which branch is this?
Which server is running?
Which workspace belongs to which task?
And which AI agent needs your attention?
cmux is designed around exactly this problem.
What isΒ cmux?
cmux is an open-source, Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents.
The project adds several developer-focused primitives:
vertical and horizontal tabs
notification rings
notification panels
integrated browser panes
Git branch and PR metadata
SSH workspaces
programmable CLI and socket APIs
custom commands
Claude Code Teams integration
And there is an important detail cmux is not trying to replace your AI agent. It provides the environment around the agent.
The project describes itself as a primitive rather than an opinionated orchestration system: terminal, browser, notifications, workspaces, splits, tabs, and a programmable interface.
That is a very interesting design philosophy.
Why Developers LoveΒ It
The problem cmux is solving becomes obvious when you run multiple coding agents.
Imagine this:
Workspace 1 β Claude Code β Authentication bug
Workspace 2 β Codex β New API endpoint
Workspace 3 β OpenCode β Refactor components
Workspace 4 β Dev server + browser
Workspace 5 β Test suiteWithout some structure, youβre constantly switching between windows.
cmux adds context directly into the terminal.
Its sidebar can show Git branch, linked PR status, working directory, listening ports and recent notification text for each workspace. When an agent needs attention, its pane can receive a visual notification ring and the corresponding tab is highlighted.
That sounds like a small UX improvement but it isnβt.
When youβre running multiple agents, knowing what needs your attention becomes a productivity problem of its own.
Real Development UseΒ Cases
Parallel Coding Agents: Run multiple AI agents against different tasks or branches.
Full-Stack Development: Keep your frontend, backend, database, and AI agents visible together.
Remote Development: Use SSH workspaces and work with remote environments.
Browser + Terminal Workflows: Open a browser next to the terminal and let agents interact with development servers through the programmable browser interface.
Team-Based Agent Workflows: The repository provides a cmux claude-teams command for Claude Code's teammate mode, creating native splits with metadata and notifications.
Workflow Automation: The CLI and socket API can create workspaces, split panes, send keystrokes and automate browser actions.
Productivity Impact
The biggest productivity gain is visibility.
You stop thinking βWhere did I put that agent?β and start thinking βWhich agent needs me right now?β
Thatβs a much better problem to have. There is one practical limitation worth remembering:
cmux is a native macOS application. The repository specifically positions it around macOS and uses Swift/AppKit with libghostty for terminal rendering.
For Mac developers working heavily with AI coding agents, though, it is a very interesting productivity primitive.
The next generation of developer tools may not replace the terminal. They may make the terminal dramatically better at working with agents.
3. Open WebUI: Your Self-Hosted AI Workspace
Repository: open-webui/open-webui
A lot of developers have experimented with local LLMs. Install Ollama, Download a model, Send a prompt then Done.
But quickly, the workflow becomes more complicated.
You want:
multiple models
RAG
tools
agents
user accounts
permissions
memory
web search
documents
analytics
automation
Suddenly, βjust running a model locallyβ isnβt enough.
Thatβs where Open WebUI becomes interesting.
What is OpenΒ WebUI?
Open WebUI is an extensible, feature-rich, self-hosted AI platform designed to operate entirely offline. It supports LLM runners such as Ollama and OpenAI-compatible APIs, with a built-in inference engine for RAG.
The project can be installed using pip, uv, Docker, or Kubernetes, and it can connect to OpenAI-compatible services as well as local models.
This means your architecture can look something like:
Open WebUI
β
βββββββββββββββΌββββββββββββββ
β β β
Ollama vLLM Cloud APIs
β β β
Local LLM Local LLM Hosted ModelsAnd then add:
RAG
Agents
MCP
Web Search
Memory
Automations
AnalyticsWhy Developers LoveΒ It
Open WebUI has grown beyond being a basic chat interface.
Its current README lists:
model and API integrations
RBAC and user groups
plugins
MCP, MCPO and OpenAPI tool integration
specialized models and agents
persistent memory
channels for team collaboration
scheduling and automations
voice and video
RAG
web search
image generation
multi-model conversations
model evaluation
PostgreSQL support
vector databases
enterprise authentication
OpenTelemetry
horizontal scaling
That is a lot and thatβs precisely why this project is useful for developers. You arenβt just getting a chat window.
Youβre getting a self-hosted AI application platform.
Real Development UseΒ Cases
Internal AI Assistant: Build a private AI interface for engineering teams.
Local AI Development: Connect Ollama or local inference engines.
RAG Applications: Upload documents and create knowledge-backed AI conversations.
Multi-Model Workflows: Compare multiple models in parallel.
Agent Applications: Wrap models with custom instructions, tools and knowledge.
Enterprise AI: Use RBAC, SSO, LDAP/Active Directory and SCIM for managed access.
AI Automation: Schedule recurring prompts and surface their output through the applicationβs calendar and chat workflows.
Productivity Impact
The real advantage is consolidation.
Instead of assembling ten different components just to experiment with an internal AI platform, developers can start with an existing open-source foundation.
And because it is self-hosted, teams have more control over:
infrastructure
data
models
access
integrations
The project even has an ecosystem around it, including Open WebUI Computer, Open Terminal, a native desktop application and knowledge-base tooling.
Open WebUI turns βI have a modelβ into βI have an AI workspace.β
That is a much bigger leap.
π‘ Enjoying this article?
Every week day, I publish practical, production-ready deep dives covering Web development, System Design, Open source projects, Tech industry trends and AI Engineering and tools.
4. Dify: From AI Prototype to Production Workflow
Repository: langgenius/dify
One of the biggest problems in AI development is that prototypes are easy but Production is not.
You can build a chatbot in an afternoon, but then reality arrives.
You need:
workflows
RAG
model management
tools
agents
observability
APIs
deployment
evaluation
That is where Dify fits.
What isΒ Dify?
Dify is an open-source LLM application development platform.
Its interface combines:
AI workflows
RAG pipelines
agent capabilities
model management
observability
application APIs
The projectβs stated goal is to help developers move from prototype to production without rebuilding the stack and that is the important part. Dify isnβt just about calling an LLM.
Itβs about building an AI application around the LLM.
Why Developers LoveΒ It
Dify uses a visual canvas for building AI workflows.
The current README lists:
visual workflows
support for hundreds of proprietary and open-source LLMs
prompt IDE
RAG pipelines
agent capabilities
50+ built-in agent tools
LLMOps
APIs for integrating Dify into your own business logic
That gives developers a much more structured application lifecycle.
Instead of Prompt β Model β Response
you can build:
Input
β
Prompt
β
Retriever
β
LLM
β
Tool
β
Decision
β
Output
β
Evaluationand the workflow itself becomes something you can iterate on.
Real Development UseΒ Cases
RAG Applications: Build document-grounded applications from ingestion through retrieval. Dify supports document extraction for PDFs, PPTs and other common formats.
AI Agents: Create agents using LLM function calling or ReAct patterns and connect built-in or custom tools.
Enterprise Assistants: Build internal knowledge assistants for teams.
Customer Support: Combine retrieval, reasoning and business tools.
Content Workflows: Create multi-step content generation and review pipelines.
AI API Backend: Use Difyβs APIs to integrate AI capabilities into an existing application.
Productivity Impact
Dify reduces the amount of glue code developers need to write before an AI application becomes useful. That doesnβt mean developers stop coding.
It means they can spend more time on:
application logic
business rules
prompt design
evaluation
integrations
and less time wiring every component together.
Self-hosting is also straightforward through Docker Compose, with the repository documenting a quick-start flow for running Dify locally.
The bigger advantage is lifecycle continuity.
You can move from experiment β workflow β API β production without throwing away the prototype.
5. Hugging Face Transformers: The Foundation Under Much of ModernΒ AI
Repository: huggingface/transformers
This final repository is different.
It isnβt primarily an application.
It isnβt a terminal.
It isnβt an AI UI.
It is closer to infrastructure for the model ecosystem itself.
And if youβve worked with modern open-source AI, chances are youβve already encountered it.
What is Transformers?
Hugging Face Transformers is a model-definition framework for state-of-the-art machine learning models covering text, computer vision, audio, video and multimodal workloads, supporting both inference and training.
The repository describes Transformers as a pivot across the ecosystem.
When a model definition is supported, it can work with many training frameworks and inference engines, including tools such as:
Axolotl
Unsloth
DeepSpeed
FSDP
PyTorch Lightning
vLLM
SGLang
llama.cpp
MLX
Thatβs why this repository matters so much.
Why Developers LoveΒ It
Transformers gives developers a common abstraction for working with modern models. Instead of learning a completely different API for every new architecture, developers can often work through familiar interfaces.
The Hugging Face Hub currently hosts 1M+ Transformers model checkpoints, according to the repository README. That makes the library one of the easiest entry points into the open model ecosystem.
And the high-level pipeline API is intentionally simple.
For example:
from transformers import pipeline
generator = pipeline(
task="text-generation",
model="Qwen/Qwen2.5-1.5B"
)
result = generator(
"The future of software development is"
)
print(result)The same high-level API supports multiple modalities and tasks, including text generation, speech recognition and image classification.
Real Development UseΒ Cases
Model Experimentation: Try different open-source models quickly.
NLP Applications: Build text classification, generation, summarization and other language workloads.
Computer Vision: Work with image-based models.
Speech: Run speech recognition and audio workloads.
Multimodal Applications: Experiment with models that combine text, vision, audio and other modalities.
Model Serving: Use Transformers model definitions alongside inference systems such as vLLM.
Fine-Tuning and Research: Use the model ecosystem alongside training frameworks and optimization libraries.
Productivity Impact
The most important contribution of Transformers is standardization. The AI model ecosystem changes incredibly quickly.
New models appear.
New architectures appear.
New modalities appear.
Without common abstractions, developers would constantly have to rebuild integrations. Transformers provides a common model-definition layer that many other parts of the ecosystem can build on.
And that creates a powerful chain:
Hugging Face Model
β
Transformers
β
ββββββββ΄βββββββββ
β β
Training Inference
β β
DeepSpeed vLLM
Axolotl SGLang
Unsloth llama.cppThatβs why Transformers isnβt just another Python package.
It is one of the connective layers of the modern open AI ecosystem.
The Bigger Picture: These Five Repositories Form an AI Engineering Stack
Look at the five repositories together. They seem unrelated, but they actually cover different parts of the same workflow.
Modern AI Engineering Stack
Hugging Face
Transformers
β
Model Layer
β
β
vLLM
Inference / Serving
β
βββββββββββββ΄ββββββββββββ
β β
Open WebUI Dify
AI Workspace AI Applications
β β
βββββββββββββ¬ββββββββββββ
β
AI Agents
β
cmux
Developer EnvironmentThat tells us something important. The future of AI development isnβt only about better models, itβs about everything surrounding the model.
The model is one layer, inference is another, applications are another, developer tooling is another and the productivity gains happen when those layers work together.
Which Repository Should You TryΒ First?
If youβre deploying open-source LLMs and need high-throughput inference, start with vLLM.
If youβre a Mac developer running several AI coding agents at the same time, try cmux.
If youβre building a self-hosted AI assistant or internal AI platform, explore Open WebUI.
If youβre building RAG pipelines, agents and AI applications, Dify is worth serious experimentation.
And if you want to understand the foundation beneath much of the open-source model ecosystem, Hugging Face Transformers is the obvious place to start.
You donβt need all five.
The point is understanding where each one fits.
Final Thoughts
Software developers primarily thought in terms of Frontend β Backend β Database
The AI stack is becoming much broader. Today, you might be thinking about Model β Inference β RAG β Agent β Tools β UI β Observability β Infrastructure
That is why open-source AI engineering has become so interesting.
Projects like Transformers make models accessible.
Projects like vLLM make those models practical to serve.
Projects like Open WebUI and Dify turn them into usable applications.
And projects like cmux change how developers interact with AI agents themselves. The important lesson isnβt that you need to learn every new AI repository.
Itβs this Understand the layer each tool solves, then compose the layers you actually need.
Thatβs how AI development moves from experimentation to engineering.
Thank You forΒ Reading!
I hope you found it helpful and informative. If you have any questions or feedback, feel free to leave a comment below. Your support and engagement mean a lot to me.
