Science World

China's Nanjing University Researchers Develop LightTok Chip That Converts Light Directly Into AI Tokens

China's Nanjing University Researchers Develop LightTok Chip That Converts Light Directly Into AI Tokens

NANJING, CHINA — Researchers from Nanjing University, in collaboration with the National University of Singapore, have developed LightTok, an ultra-low-power vision sensor chip that converts light directly into AI-ready tokens inside the sensor. The findings were published this week in the journal Nature Sensors.

 

Direct AI Tokenization at the Sensor

Conventional vision systems capture light, convert analog signals into digital data, buffer and transfer the data, and then divide images into patches and embed them before producing tokens for Transformer models. These steps involve moving large amounts of often redundant data and consume significant energy, particularly in power-constrained edge devices.

LightTok performs this process directly in the physical domain. Its photosensitive memory array uses single-layer molybdenum disulfide (MoS₂) floating-gate phototransistors along with peripheral addressing circuits.

Each pixel can sense light, store information as non-volatile charge and perform analog computing. Peripheral circuits select regions of the array and divide the image into blocks. Voltage sequences then allow the stored optical information to perform multiply-add operations in the analog domain, with the resulting current becoming the output token.

These tokens can be sent directly to a Transformer encoder for image recognition. Liang Shijun, a professor at Nanjing University, described the approach as eliminating data movement that contributes significantly to energy consumption, with light entering the sensor and tokens being produced directly.

 

14.3-Fold Energy Reduction

In tests, LightTok achieved 87.3 percent accuracy on image-recognition tasks, close to conventional software baselines.

The research reported that energy efficiency during tokenization improved by more than 10 times compared with traditional digital methods. In a standard Vision Transformer pipeline using the CIFAR-10 dataset, the physical tokenizer achieved a 14.3-fold reduction in energy use compared with a digital tokenizer.

 

32 × 32-Pixel Prototype

The current prototype has a 32 × 32-pixel resolution, providing 1,024 photosensitive pixels. This is smaller than sensors used in smartphones and industrial cameras.

Miao Feng, director of Nanjing University's Institute of Brain-Inspired Intelligence and a co-leader of the research, said the technology is compatible with standard CMOS manufacturing processes and can be scaled. The researchers said wafer-level growth of molybdenum disulfide could eventually allow LightTok to reach resolutions comparable to existing imaging devices.

The design is inspired by the human retina, which performs early visual processing before sending information to the brain.

 

Potential Applications

The researchers identified potential uses in drones, autonomous satellites for remote sensing, and small embodied-AI systems that require continuous visual processing under tight power constraints.

By generating AI-ready tokens directly at the sensing stage, LightTok is intended to reduce the energy required to acquire and tokenize visual information for edge AI and physical AI systems.

——— End of Article ———

About the Author

Aditya Kumar is a Defense & Geopolitics Analyst covering military developments, missile systems, naval strategy, and global defense affairs.