Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning

About

Natural language processing (NLP) has significantly influenced scientific domains beyond human language, including protein engineering, where pre-trained protein language models (PLMs) have demonstrated remarkable success. However, interdisciplinary adoption remains limited due to challenges in data collection, task benchmarking, and application. This work presents VenusFactory, a versatile engine that integrates biological data retrieval, standardized task benchmarking, and modular fine-tuning of PLMs. VenusFactory supports both computer science and biology communities with choices of both a command-line execution and a Gradio-based no-code interface, integrating $40+$ protein-related datasets and $40+$ popular PLMs. All implementations are open-sourced on https://github.com/tyang816/VenusFactory.

Yang Tan, Chen Liu, Jingyuan Gao, Banghao Wu, Mingchen Li, Ruilin Wang, Lingrong Zhang, Huiqun Yu, Guisheng Fan, Liang Hong, Bingxin Zhou• 2025

Related benchmarks

TaskDatasetResultRank
Protein property predictionPFMBench
Antibiotic Resistance Accuracy64.6
13
Showing 1 of 1 rows

Other info

Follow for update