torchchat: Run PyTorch LLMs on Desktop, Server, and Mobile

torchchat: Run PyTorch LLMs on Desktop, Server, and Mobile

Summary

torchchat is a PyTorch codebase for running and interacting with language models locally through Python, native C++ runners, and mobile apps. It supports several execution and export paths, but is no longer under active development.

At a glance

Language
Python
License
BSD-3-Clause
Stars
3.6k
Forks
248
Added to OSRepos
July 3, 2026
Last analyzed
October 3, 2026

This repository is archived on GitHub and no longer maintained.

View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

torchchat provides a way to run supported large language models with PyTorch across Python environments, desktop and server systems, and mobile devices. It brings model downloads, chat and text generation, export, and evaluation into one toolkit, addressing the challenge of adapting local inference to different execution environments.

It is intended for developers experimenting with or integrating local LLM inference, especially where PyTorch-native execution or deployment through AOT Inductor and ExecuTorch is useful. The project is no longer under active development, so it is best treated as an existing toolkit rather than a growing platform.

Key Features

  • Run interactive chat and prompt-based text generation from Python.
  • Serve chat completions through a locally hosted REST API with an OpenAI-style interface.
  • Launch a basic browser chat interface against a local server.
  • Export models for AOT Inductor or ExecuTorch, with Python and C++ runner options.
  • Deploy exported models to iOS and Android using ExecuTorch.
  • Use supported model aliases spanning Llama, Mistral, Granite, DeepSeek, and other listed families.
  • Configure data types and quantization, including mobile-oriented quantization options.
  • Evaluate models through an integration with lm-evaluation-harness.

Use Cases

  • A PyTorch developer wants to test local chat or text generation without building an inference stack from scratch.
  • An application developer wants to export a supported model and run it from a C++ desktop or server process.
  • A mobile developer wants to explore on-device LLM inference using ExecuTorch and an exported model artifact.
  • A team comparing model outputs or configurations wants to run the repository's evaluation workflow on a supported model.

Project Facts

  • Language: Python
  • License: BSD-3-Clause
  • Stars: 3.6k
  • Forks: 248
  • Topics: llm, local, pytorch
  • Archived: true

Getting Started

The README specifies Python 3.10 and recommends using a virtual environment. A minimal setup is:

git clone https://github.com/pytorch/torchchat.git
cd torchchat
python3 -m venv .venv
source .venv/bin/activate
./install/install_requirements.sh

See the README for model access, usage commands, and platform-specific deployment instructions.

Alternatives

  • cactus: Cactus is a C++-focused engine for local language, vision, and speech inference, while torchchat offers PyTorch, C++, and mobile execution paths for language models.

Considerations

  • The repository states that it is no longer under active development, and it is archived. Expect limited maintenance and no assumption of ongoing fixes.
  • The README recommends Python 3.10 and a virtual environment. Mobile deployment also requires setting up ExecuTorch and the relevant iOS or Android development tools.
  • Many model weights are distributed through Hugging Face. Some models require account access or additional authorization, and third-party model terms may apply.
  • The REST server and evaluation workflow are identified in the README as works in progress, with some parameters or features not fully implemented.
  • Model size, hardware compatibility, and performance vary. The README does not guarantee performance or compatibility.

Source repository

Open the original repository on GitHub.

24 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design

web-design: A Claude Code SKILL for Spec-First Web Page Design

October 3, 2026

The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

Claude CodeClaude SkillDesign System
OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner

October 2, 2026

OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

RoboticsOpen SourceDiy
Shepherd: Reversible Execution Traces for Programmable Meta-Agents

Shepherd: Reversible Execution Traces for Programmable Meta-Agents

October 2, 2026

Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

PythonAIAgent Framework
Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

October 1, 2026

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

PythonAIAgent
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️