Most AI developers were focused on one thing Getting an LLM to work. But, that is no longer enough.

The real engineering challenge is increasingly about:

  • Serving models efficiently

  • Running multiple AI agents at the same time

  • Building self-hosted AI interfaces

  • Creating reliable AI applications

  • Connecting models to production workflows

  • Supporting hundreds of models without rewriting your stack

In other words, We are moving from AI experimentation to AI infrastructure.

For Week 19 of my Open Source GitHub Repository Series, I explored five repositories that represent very different layers of that infrastructure:

  • vLLM β†’ High-throughput LLM inference and serving

  • cmux β†’ A developer terminal designed around parallel AI coding agents

  • Open WebUI β†’ A self-hosted AI platform for models, agents, RAG, tools, and workflows

  • Dify β†’ An open-source platform for building agentic workflows and AI applications

  • Hugging Face Transformers β†’ The model-definition framework powering a huge portion of the modern AI ecosystem

What makes this week’s list especially interesting is that these projects operate at different levels of the stack.

  • Some help you run the model.

  • Some help you work with the model.

  • Some help you build applications around the model.

And one helps connect the broader model ecosystem together.

Let’s dive in.

1. vLLM: The Inference Engine Behind Production-Scale AI

Repository: vllm-project/vllm

There is a big difference between running a model and serving a model efficiently.

Downloading an LLM and getting one response is relatively easy. The challenge begins when thousands of users start sending requests.

Now you care about:

  • throughput

  • latency

  • memory usage

  • batching

  • GPU utilization

  • concurrency

  • distributed inference

  • model compatibility

That is where vLLM becomes important.

What isΒ vLLM?

vLLM is an open-source library for fast and easy-to-use LLM inference and serving.

The project focuses heavily on throughput and memory efficiency and includes techniques such as PagedAttention, continuous batching, chunked prefill, prefix caching, optimized attention kernels, quantization, speculative decoding, and distributed parallelism.

The project also provides an OpenAI-compatible API server, along with Anthropic Messages API and gRPC support.

That makes vLLM much more than a research library.

It can become the serving layer sitting between your application and the model.

Think of the architecture as:

Your Application
       ↓
     vLLM
       ↓
    LLM Model
       ↓
    GPU / TPU

Instead of every application figuring out how to efficiently execute models, vLLM handles the serving layer.

Why Developers LoveΒ It

One of the hardest parts of LLM infrastructure is making expensive hardware useful. A GPU that spends most of its time waiting for requests is wasted infrastructure.

vLLM focuses directly on increasing serving efficiency.

Its README highlights state-of-the-art serving throughput, efficient KV-cache memory management, continuous batching, prefix caching, quantization support, optimized kernels, speculative decoding, and distributed inference capabilities.

It also supports a broad range of model architectures and hardware.

The repository currently describes support for 200+ model architectures on Hugging Face, including decoder-only LLMs, Mixture-of-Experts models, hybrid architectures, multimodal models, embedding/retrieval models, and reward/classification models. It also lists support across NVIDIA, AMD, Intel and other hardware ecosystems.

That makes it especially interesting for teams building their own AI infrastructure.

Real Development UseΒ Cases

Self-Hosted LLM APIs: Run an open-source model behind an API that your applications can consume.

AI SaaS Platforms: Serve models for multiple customers while managing concurrency and infrastructure efficiently.

Internal AI Platforms: Provide a common model-serving layer for multiple internal applications.

RAG Systems: Deploy the generation model behind a retrieval pipeline.

Agent Platforms: Serve the models used by coding agents and autonomous workflows.

Model Experimentation: Quickly switch between supported models without rebuilding the whole serving layer.

Productivity Impact

Developers don’t want to spend their time writing low-level inference infrastructure.

They want to build products.

vLLM provides a mature serving layer so teams can focus more on:

  • application logic

  • agent behavior

  • evaluation

  • product features

and less on reinventing model-serving infrastructure.

The important distinction is this: Hugging Face Transformers helps define and use models. vLLM helps serve those models efficiently in real systems.

That difference becomes extremely important once your AI application leaves the prototype stage.

2. cmux: A Terminal Built for Developers Running AI Agents inΒ Parallel

Repository: manaflow-ai/cmux

Modern coding workflows are becoming strangely crowded.

You might have:

  • Claude Code

  • Codex

  • OpenCode

  • a local terminal

  • a development server

  • browser tools

  • Git

  • multiple repositories

  • several tasks running simultaneously

Open ten terminal tabs and eventually the terminal becomes the problem.

  • Which agent is waiting?

  • Which branch is this?

  • Which server is running?

  • Which workspace belongs to which task?

And which AI agent needs your attention?

cmux is designed around exactly this problem.

What isΒ cmux?

cmux is an open-source, Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents.

The project adds several developer-focused primitives:

  • vertical and horizontal tabs

  • notification rings

  • notification panels

  • integrated browser panes

  • Git branch and PR metadata

  • SSH workspaces

  • programmable CLI and socket APIs

  • custom commands

  • Claude Code Teams integration

And there is an important detail cmux is not trying to replace your AI agent. It provides the environment around the agent.

The project describes itself as a primitive rather than an opinionated orchestration system: terminal, browser, notifications, workspaces, splits, tabs, and a programmable interface.

That is a very interesting design philosophy.

Why Developers LoveΒ It

The problem cmux is solving becomes obvious when you run multiple coding agents.

Imagine this:

Workspace 1 β†’ Claude Code β†’ Authentication bug
Workspace 2 β†’ Codex β†’ New API endpoint
Workspace 3 β†’ OpenCode β†’ Refactor components
Workspace 4 β†’ Dev server + browser
Workspace 5 β†’ Test suite

Without some structure, you’re constantly switching between windows.

cmux adds context directly into the terminal.

Its sidebar can show Git branch, linked PR status, working directory, listening ports and recent notification text for each workspace. When an agent needs attention, its pane can receive a visual notification ring and the corresponding tab is highlighted.

That sounds like a small UX improvement but it isn’t.

When you’re running multiple agents, knowing what needs your attention becomes a productivity problem of its own.

Real Development UseΒ Cases

Parallel Coding Agents: Run multiple AI agents against different tasks or branches.

Full-Stack Development: Keep your frontend, backend, database, and AI agents visible together.

Remote Development: Use SSH workspaces and work with remote environments.

Browser + Terminal Workflows: Open a browser next to the terminal and let agents interact with development servers through the programmable browser interface.

Team-Based Agent Workflows: The repository provides a cmux claude-teams command for Claude Code's teammate mode, creating native splits with metadata and notifications.

Workflow Automation: The CLI and socket API can create workspaces, split panes, send keystrokes and automate browser actions.

Productivity Impact

The biggest productivity gain is visibility.

You stop thinking β€œWhere did I put that agent?” and start thinking β€œWhich agent needs me right now?”

That’s a much better problem to have. There is one practical limitation worth remembering:

cmux is a native macOS application. The repository specifically positions it around macOS and uses Swift/AppKit with libghostty for terminal rendering.

For Mac developers working heavily with AI coding agents, though, it is a very interesting productivity primitive.

The next generation of developer tools may not replace the terminal. They may make the terminal dramatically better at working with agents.

3. Open WebUI: Your Self-Hosted AI Workspace

A lot of developers have experimented with local LLMs. Install Ollama, Download a model, Send a prompt then Done.

But quickly, the workflow becomes more complicated.

You want:

  • multiple models

  • RAG

  • tools

  • agents

  • user accounts

  • permissions

  • memory

  • web search

  • documents

  • analytics

  • automation

Suddenly, β€œjust running a model locally” isn’t enough.

That’s where Open WebUI becomes interesting.

What is OpenΒ WebUI?

Open WebUI is an extensible, feature-rich, self-hosted AI platform designed to operate entirely offline. It supports LLM runners such as Ollama and OpenAI-compatible APIs, with a built-in inference engine for RAG.

The project can be installed using pip, uv, Docker, or Kubernetes, and it can connect to OpenAI-compatible services as well as local models.

This means your architecture can look something like:

            Open WebUI
                 β”‚
   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
   β”‚             β”‚             β”‚
Ollama        vLLM        Cloud APIs
   β”‚             β”‚             β”‚
Local LLM     Local LLM    Hosted Models

And then add:

RAG
Agents
MCP
Web Search
Memory
Automations
Analytics

Why Developers LoveΒ It

Open WebUI has grown beyond being a basic chat interface.

Its current README lists:

  • model and API integrations

  • RBAC and user groups

  • plugins

  • MCP, MCPO and OpenAPI tool integration

  • specialized models and agents

  • persistent memory

  • channels for team collaboration

  • scheduling and automations

  • voice and video

  • RAG

  • web search

  • image generation

  • multi-model conversations

  • model evaluation

  • PostgreSQL support

  • vector databases

  • enterprise authentication

  • OpenTelemetry

  • horizontal scaling

That is a lot and that’s precisely why this project is useful for developers. You aren’t just getting a chat window.

You’re getting a self-hosted AI application platform.

Real Development UseΒ Cases

Internal AI Assistant: Build a private AI interface for engineering teams.

Local AI Development: Connect Ollama or local inference engines.

RAG Applications: Upload documents and create knowledge-backed AI conversations.

Multi-Model Workflows: Compare multiple models in parallel.

Agent Applications: Wrap models with custom instructions, tools and knowledge.

Enterprise AI: Use RBAC, SSO, LDAP/Active Directory and SCIM for managed access.

AI Automation: Schedule recurring prompts and surface their output through the application’s calendar and chat workflows.

Productivity Impact

The real advantage is consolidation.

Instead of assembling ten different components just to experiment with an internal AI platform, developers can start with an existing open-source foundation.

And because it is self-hosted, teams have more control over:

  • infrastructure

  • data

  • models

  • access

  • integrations

The project even has an ecosystem around it, including Open WebUI Computer, Open Terminal, a native desktop application and knowledge-base tooling.

Open WebUI turns β€œI have a model” into β€œI have an AI workspace.”

That is a much bigger leap.

πŸ’‘ Enjoying this article?
Every week day, I publish practical, production-ready deep dives covering Web development, System Design, Open source projects, Tech industry trends and AI Engineering and tools.

4. Dify: From AI Prototype to Production Workflow

Repository: langgenius/dify

One of the biggest problems in AI development is that prototypes are easy but Production is not.

You can build a chatbot in an afternoon, but then reality arrives.

You need:

  • workflows

  • RAG

  • model management

  • tools

  • agents

  • observability

  • APIs

  • deployment

  • evaluation

That is where Dify fits.

What isΒ Dify?

Dify is an open-source LLM application development platform.

Its interface combines:

  • AI workflows

  • RAG pipelines

  • agent capabilities

  • model management

  • observability

  • application APIs

The project’s stated goal is to help developers move from prototype to production without rebuilding the stack and that is the important part. Dify isn’t just about calling an LLM.

It’s about building an AI application around the LLM.

Why Developers LoveΒ It

Dify uses a visual canvas for building AI workflows.

The current README lists:

  • visual workflows

  • support for hundreds of proprietary and open-source LLMs

  • prompt IDE

  • RAG pipelines

  • agent capabilities

  • 50+ built-in agent tools

  • LLMOps

  • APIs for integrating Dify into your own business logic

That gives developers a much more structured application lifecycle.

Instead of Prompt β†’ Model β†’ Response

you can build:

Input
  ↓
Prompt
  ↓
Retriever
  ↓
LLM
  ↓
Tool
  ↓
Decision
  ↓
Output
  ↓
Evaluation

and the workflow itself becomes something you can iterate on.

Real Development UseΒ Cases

RAG Applications: Build document-grounded applications from ingestion through retrieval. Dify supports document extraction for PDFs, PPTs and other common formats.

AI Agents: Create agents using LLM function calling or ReAct patterns and connect built-in or custom tools.

Enterprise Assistants: Build internal knowledge assistants for teams.

Customer Support: Combine retrieval, reasoning and business tools.

Content Workflows: Create multi-step content generation and review pipelines.

AI API Backend: Use Dify’s APIs to integrate AI capabilities into an existing application.

Productivity Impact

Dify reduces the amount of glue code developers need to write before an AI application becomes useful. That doesn’t mean developers stop coding.

It means they can spend more time on:

  • application logic

  • business rules

  • prompt design

  • evaluation

  • integrations

and less time wiring every component together.

Self-hosting is also straightforward through Docker Compose, with the repository documenting a quick-start flow for running Dify locally.

The bigger advantage is lifecycle continuity.

You can move from experiment β†’ workflow β†’ API β†’ production without throwing away the prototype.

5. Hugging Face Transformers: The Foundation Under Much of ModernΒ AI

This final repository is different.

  • It isn’t primarily an application.

  • It isn’t a terminal.

  • It isn’t an AI UI.

  • It is closer to infrastructure for the model ecosystem itself.

And if you’ve worked with modern open-source AI, chances are you’ve already encountered it.

What is Transformers?

Hugging Face Transformers is a model-definition framework for state-of-the-art machine learning models covering text, computer vision, audio, video and multimodal workloads, supporting both inference and training.

The repository describes Transformers as a pivot across the ecosystem.

When a model definition is supported, it can work with many training frameworks and inference engines, including tools such as:

  • Axolotl

  • Unsloth

  • DeepSpeed

  • FSDP

  • PyTorch Lightning

  • vLLM

  • SGLang

  • llama.cpp

  • MLX

That’s why this repository matters so much.

Why Developers LoveΒ It

Transformers gives developers a common abstraction for working with modern models. Instead of learning a completely different API for every new architecture, developers can often work through familiar interfaces.

The Hugging Face Hub currently hosts 1M+ Transformers model checkpoints, according to the repository README. That makes the library one of the easiest entry points into the open model ecosystem.

And the high-level pipeline API is intentionally simple.

For example:

from transformers import pipeline

generator = pipeline(
    task="text-generation",
    model="Qwen/Qwen2.5-1.5B"
)
result = generator(
    "The future of software development is"
)
print(result)

The same high-level API supports multiple modalities and tasks, including text generation, speech recognition and image classification.

Real Development UseΒ Cases

Model Experimentation: Try different open-source models quickly.

NLP Applications: Build text classification, generation, summarization and other language workloads.

Computer Vision: Work with image-based models.

Speech: Run speech recognition and audio workloads.

Multimodal Applications: Experiment with models that combine text, vision, audio and other modalities.

Model Serving: Use Transformers model definitions alongside inference systems such as vLLM.

Fine-Tuning and Research: Use the model ecosystem alongside training frameworks and optimization libraries.

Productivity Impact

The most important contribution of Transformers is standardization. The AI model ecosystem changes incredibly quickly.

  • New models appear.

  • New architectures appear.

  • New modalities appear.

Without common abstractions, developers would constantly have to rebuild integrations. Transformers provides a common model-definition layer that many other parts of the ecosystem can build on.

And that creates a powerful chain:

Hugging Face Model
        ↓
   Transformers
        ↓
 β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
 ↓               ↓
Training       Inference
 ↓               ↓
DeepSpeed      vLLM
Axolotl        SGLang
Unsloth        llama.cpp

That’s why Transformers isn’t just another Python package.

It is one of the connective layers of the modern open AI ecosystem.

The Bigger Picture: These Five Repositories Form an AI Engineering Stack

Look at the five repositories together. They seem unrelated, but they actually cover different parts of the same workflow.

Modern AI Engineering Stack

           Hugging Face
           Transformers
                β”‚
           Model Layer
                β”‚
                ↓
             vLLM
        Inference / Serving
                β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚                       β”‚
 Open WebUI                Dify
 AI Workspace          AI Applications
    β”‚                       β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                β”‚
            AI Agents
                β”‚
               cmux
      Developer Environment

That tells us something important. The future of AI development isn’t only about better models, it’s about everything surrounding the model.

The model is one layer, inference is another, applications are another, developer tooling is another and the productivity gains happen when those layers work together.

Which Repository Should You TryΒ First?

If you’re deploying open-source LLMs and need high-throughput inference, start with vLLM.

If you’re a Mac developer running several AI coding agents at the same time, try cmux.

If you’re building a self-hosted AI assistant or internal AI platform, explore Open WebUI.

If you’re building RAG pipelines, agents and AI applications, Dify is worth serious experimentation.

And if you want to understand the foundation beneath much of the open-source model ecosystem, Hugging Face Transformers is the obvious place to start.

You don’t need all five.

The point is understanding where each one fits.

Final Thoughts

Software developers primarily thought in terms of Frontend β†’ Backend β†’ Database

The AI stack is becoming much broader. Today, you might be thinking about Model β†’ Inference β†’ RAG β†’ Agent β†’ Tools β†’ UI β†’ Observability β†’ Infrastructure

That is why open-source AI engineering has become so interesting.

  • Projects like Transformers make models accessible.

  • Projects like vLLM make those models practical to serve.

  • Projects like Open WebUI and Dify turn them into usable applications.

And projects like cmux change how developers interact with AI agents themselves. The important lesson isn’t that you need to learn every new AI repository.

It’s this Understand the layer each tool solves, then compose the layers you actually need.

That’s how AI development moves from experimentation to engineering.

Thank You forΒ Reading!

I hope you found it helpful and informative. If you have any questions or feedback, feel free to leave a comment below. Your support and engagement mean a lot to me.

Happy Coding!

Reply

Avatar

or to participate