# SpatialLM: Turn 3D Point Clouds into Structured Indoor Scenes

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Source: osrepos.com
Repository profile: https://osrepos.com/repo/manycore-research-spatiallm
Generated for open source discovery and AI-assisted research.

SpatialLM uses language models to interpret indoor point clouds and produce structured layouts, including walls, doors, windows, and object boxes. It suits 3D perception researchers and teams building scene-understanding prototypes, but requires a CUDA-oriented setup.

GitHub: https://github.com/manycore-research/SpatialLM
OSRepos URL: https://osrepos.com/repo/manycore-research-spatiallm

## Summary

SpatialLM uses language models to interpret indoor point clouds and produce structured layouts, including walls, doors, windows, and object boxes. It suits 3D perception researchers and teams building scene-understanding prototypes, but requires a CUDA-oriented setup.

## Topics

- python
- llm
- computer-vision
- 3d-reconstruction
- point-clouds
- indoor-scene-understanding

## Repository Information

Last analyzed by OSRepos: Sat Oct 10 2026 21:07:34 GMT+0100 (Western European Summer Time)
Detail views: 1
GitHub clicks: 0

## Safety Notice

OSRepos shares public repositories for knowledge and discovery only. Review source code, dependencies, licenses, and security implications before running or installing anything.

## Content

## Overview

SpatialLM is a 3D language-model project for converting indoor point clouds into structured scene descriptions. It predicts architectural elements such as walls, doors, and windows, as well as oriented boxes with semantic categories, bridging raw geometry and representations that downstream spatial reasoning can use.

The project is aimed at researchers and developers working on indoor scene understanding, robotics, or navigation. It supports point clouds originating from sources such as monocular video reconstruction, RGB-D images, and LiDAR, provided they are oriented with the z-axis up.

## Key Features

- Predicts indoor architectural layouts, including walls, doors, and windows.
- Estimates oriented 3D object boxes and semantic categories.
- Offers separate structured reconstruction, layout estimation, and object detection task modes.
- SpatialLM1.1 can condition object detection on user-specified categories.
- Provides Llama- and Qwen-based model variants, with model weights hosted on Hugging Face.
- Includes inference, evaluation, and Rerun visualization scripts.
- Publishes a synthetic training dataset and a test set of point clouds reconstructed from RGB video.
- Documents fine-tuning on custom data, with an ARKitScenes example.

## Use Cases

- **Indoor scene-understanding research:** Researchers can test structured layout or object predictions against point-cloud benchmarks.
- **Video-to-layout prototypes:** 3D vision teams can reconstruct a point cloud from RGB video and use SpatialLM to estimate walls, doors, windows, and furniture.
- **Embodied robotics and navigation:** Developers can explore using semantic scene structure as input to spatial reasoning or navigation systems.
- **Category-focused detection:** Teams evaluating a limited set of furniture classes can specify categories for SpatialLM1.1 object detection.

## Our Take

SpatialLM is a useful research toolkit with published models, datasets, and evaluation scripts, but its quiet maintenance and concentrated contributor base call for careful adoption.

**Good fit if:**
- You work on indoor 3D perception and can supply axis-aligned point clouds with z as the up axis.
- You want to evaluate or fine-tune language-model-based layout and object detection using the provided scripts and datasets.
- Your team can manage the CUDA-oriented Python setup and review the applicable code and model licenses.

**Look elsewhere if:**
- You need evidence of active ongoing development: there have been no commits in the last 90 days and no release in over a year.
- You require a project maintained by a broad contributor base. Three contributors are recorded, and most contributions come from one person.
- You need a clearly uniform license across the project without separately reviewing its derived models and components.

## Project Health

| Signal | Value |
|---|---|
| Status | **Quiet**: no commits in the last 90 days and no release in a year |
| Last commit | 2026-06-26 (4 months ago) |
| Commits, last 90 days | 0 |
| Releases, last 12 months | 0 (latest v0.1.1, 2025-06-10) |
| Contributors | 3 (top contributor: 90% of commits) |
| Issues closed, last 90 days | 3 (typically closed in 15 days) |
| Pull requests merged, last 90 days | 0 |
| Project age | 19 months |

Checked on 2026-10-10 with the GitHub API.

## Project Facts

- Language: Python
- Stars: 4.9k
- Forks: 409
- Topics: mllm, point-clouds, scene-understanding, spatial-intelligence
- Archived: no

## What You Need

Detected in the repository:

- Python (from pyproject.toml)

## Getting Started

The documented setup uses Python 3.11, PyTorch 2.4.1, and CUDA 12.4. After installing the dependencies described in the [README](https://github.com/manycore-research/SpatialLM#installation), run inference with a point cloud and a hosted model:

```bash
python inference.py --point_cloud pcd/scene0000_00.ply --output scene0000_00.txt --model_path manycore-research/SpatialLM1.1-Qwen-0.5B
```

See the [repository README](https://github.com/manycore-research/SpatialLM) for installation, evaluation, visualization, and fine-tuning instructions.

## License in Plain Words

No standard license was detected. Without a license, the default copyright rules apply: you may read the code, but not reuse it. Check the repository or ask the authors.

## Considerations

- The measured project status is Quiet: there have been no commits in the last 90 days and no release in a year. Three contributors are recorded, with roughly 90% of contributions concentrated in one.
- Installation is more involved than a simple Python package: the documented environment includes CUDA 12.4 and builds TorchSparse or flash-attn-related dependencies depending on the model generation.
- Input point clouds are expected to be axis-aligned with z as the up axis, so other coordinate conventions may require preprocessing.
- The repository API reports `NOASSERTION` for the license. The README describes different licenses for derived Llama and Qwen models, point-cloud encoders, and code components. Check the applicable terms for the exact model and dependencies before deployment.
- The README documents research tasks and fine-tuning, but does not establish production support or a general-purpose scene representation beyond the described structured outputs.