AI Tech Processing API: Async FastAPI service wrapping local Ollama (Qwen 3) for text processing

An asynchronous AI-powered text processing REST API built with FastAPI and local LLM inference through Ollama. The service provides summarization, text classification, named entity extraction, rewriting, and translation through validated API endpoints. The project demonstrates a production-minded AI Engineering architecture by separating API routing, validation, AI services, and model providers. It includes asynchronous model communication, structured JSON generation, Pydantic validation, retry mechanisms with exponential backoff, structured error handling, and containerized deployment with Docker. The system uses Qwen 3 1.7B as the default local inference model, allowing text processing to run locally without relying on external AI APIs.

Python 3.12FastAPIPydantic v2UvicornhttpxOllama
AI Tech Processing API: Async FastAPI service wrapping local Ollama (Qwen 3) for text processing
2026

Key Features

Powerful capabilities that drive innovation

Text Summarization
Condenses long-form text into concise summaries through a dedicated REST API endpoint.
Text Classification
Classifies text by category and sentiment using schema-constrained structured AI output.
Named Entity Extraction
Identifies and extracts named entities such as people
AI Text Rewriting
organizations
Language Translation
and locations into structured JSON.
Local LLM Inference
Rewrites input text into a specified tone while preserving its original meaning.

Technology Stack

Modern technologies powering the platform

Python 3.12
Python 3.12 technology
FastAPI
FastAPI technology
Pydantic v2
Pydantic v2 technology
Uvicorn
Uvicorn technology
httpx
httpx technology
Ollama
Ollama technology
Qwen 3 1.7B
Qwen 3 1.7B technology
Tenacity
Tenacity technology
Docker
Docker technology
Docker Compose
Docker Compose technology
pytest
pytest technology

Ready to Explore?

Experience the innovation and technical excellence of this project