An asynchronous AI-powered text processing REST API built with FastAPI and local LLM inference through Ollama. The service provides summarization, text classification, named entity extraction, rewriting, and translation through validated API endpoints. The project demonstrates a production-minded AI Engineering architecture by separating API routing, validation, AI services, and model providers. It includes asynchronous model communication, structured JSON generation, Pydantic validation, retry mechanisms with exponential backoff, structured error handling, and containerized deployment with Docker. The system uses Qwen 3 1.7B as the default local inference model, allowing text processing to run locally without relying on external AI APIs.
Powerful capabilities that drive innovation
Modern technologies powering the platform
Experience the innovation and technical excellence of this project