TEM-LLM
TEM-LLM is a domain-specific multimodal large language model for Transmission Electron Microscopy (TEM) image analysis, fine-tuned from LLaVA-v1.5-7B using a difficulty-aware curriculum learning strategy.
This model is associated with the following work (currently under review):
AI-Assisted Materials Characterization: A Curriculum-Guided Multimodal Framework for Transmission Electron Microscopy
Model Details
| Base Model | LLaVA-v1.5-7B (Vicuna-v1.5-7B) |
| Training Strategy | Difficulty-Aware Curriculum Learning (4 stages) |
| Training Data | ~216K QA pairs across 40K TEM images |
| Domain | Transmission Electron Microscopy (TEM) |
| Modalities | CTEM, HR-TEM, STEM, Diffraction |
| Fine-tuning Method | LoRA (Low-Rank Adaptation) |
| License | LLaMA 2 |
Intended Use
TEM-LLM is designed for:
- Image Classification: Identifying TEM imaging modalities (CTEM, HRTEM, STEM, diffraction)
- Microstructural Feature Detection: Localizing and counting structural elements such as nanoparticles and grain boundaries
- Feature Analysis & Interpretation: Providing scientific explanations of observed microstructural phenomena
- Comprehensive Image Description: Integrating multiple aspects of visual understanding into structured descriptions
Important: Inference Requirements
TEM-LLM is built on LLaVA and cannot be loaded directly via
transformers. Inference requires the LLaVA repository.
Step 1 โ Clone LLaVA:
git clone https://github.com/haotian-liu/LLaVA.git
cd LLaVA
pip install -e .
Step 2 โ Download weights:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="LabSmart/TEM-LLM",
local_dir="./TEM-LLM"
)
Step 3 โ Run inference:
python -m llava.serve.cli \
--model-path "./TEM-LLM" \
--image-file "path/to/your/tem_image.jpg" \
--load-4bit
For the complete pipeline including data crawling, preprocessing, knowledge distillation, training, and evaluation scripts, please refer to the GitHub repository.
Limitations
- TEM-LLM is optimized for TEM image analysis and may produce unreliable outputs when applied to out-of-scope questions or out-of-domain images.
- The model provides approximate qualitative descriptions rather than exact quantitative measurements.
- Downloads last month
- 5
Model tree for LabSmart/TEM-LLM
Base model
liuhaotian/llava-v1.5-7b