Extracting AI agent-accessible data from biodiversity literature with corpus.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
Much of our shared knowledge about biodiversity, including data on species' traits, relationships, and occurrences, is stored in the published literature. This body of knowledge can span multiple centuries, languages, and publication formats, making it challenging to collect, compare, and synthesize data across entire clades. New artificial intelligence (AI) tools, including large language models (LLMs) in particular, are well suited to integrate and summarize large bodies of text and figures. However, the biodiversity literature is too large and multi-layered to be reliably interrogated by simply placing publications into an LLM context window. For the typical tasks of biodiversity scientists (e.g., assembling descriptive characters, reconstructing nomenclatural history), we need structured, provenance-aware, and deterministic data retrieval. Here we present corpus: a tool for converting biodiversity literature PDFs for a clade of organisms into a searchable knowledge base (a 'corpuscle') of text, figures, and references. Users connect an LLM client via corpus' Model Context Protocol (MCP) server, to interact directly with the processed literature through predefined search tools. corpus addresses critical gaps in LLM applications for biology by transforming the literature into an explicit retrieval layer, rather than relying on hidden training data or probabilistic recollection.
متن کامل اصلی
برای بررسی دسترسی کتابخانهای یا خرید، رکورد اصلی را باز کنید.
رفتن به منبع اصلی