Projects with this topic
-
A self-hosted AI platform — inference, tool use, browser automation, image generation, speech synthesis, transcription, object storage, agentic code execution, and more — behind a single OpenAI-compatible endpoint. One docker-compose up.
Updated -
High-performance Ultralytics YOLO inference in Rust with ONNX Runtime, GPU backends, CLI, and WebGPU/WASM.
Updated -
-
Ready-to-use Cog deployments and CI/CD for running Ultralytics YOLO11, YOLO World, YOLOE, and YOLO26 models on Replicate.
Updated -
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Updated -
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
Updated -
Docker container, inference pipeline, and submission tooling for deploying trained YOLOv3 object detection models in the xView satellite imagery challenge.
Updated -
YOLOv3 training, preprocessing, validation, and inference for object detection in xView satellite imagery and the xView detection challenge.
Updated -
Fast Bayesian authorship attribution software based on Poisson-Dirichlet processes.
Updated -
POC: ONNX Runtime inference inside a Raylib real-time graphics loop
Updated -
-
A powerful, single-file frontend for LM Studio — run local LLMs with a professional chat interface, knowledge base, guardrails, and full document support. No installation. No server. Just open the HTML file.
Updated -
Hive is a peer-to-peer system that distributes AI inference tasks across volunteer workers ("bees") running local or cloud LLMs. Send the same task to multiple bees in parallel, then automatically merge their outputs over several rounds to make small models smarter together. Built with Rust, Tauri, and libp2p.
Updated -
Intelligent VRAM/RAM swapping for LLM inference - Extension of KVortex | Offloading intelligent VRAM/RAM pour l'inference
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated -
VRAM to RAM Offloader for AI and vLLM - High-Performance C++23 KV Cache Engine with Multi-Stream GPU Transfers
Updated -
Extreme KV Cache Compression for LLM Inference — C++17/CUDA implementation of TurboQuant (arXiv 2504.19874). 7.5x compression, <2% quality loss.
Updated -
Evaluation of Fast, Faster and Mask R-CNN regarding their inference times
Updated