Pharma · Chemical R&D · Caso de estudio
Instead of toggling between three windows — reports, patents, papers — you ask from one channel, and every answer arrives with a document name, a page, metadata
Lo que consiguesAsk internal technical reports, patents, and external papers from one channel without moving windows, and get a report where the document name and page, the patent conditions, and the paper metadata are attached as sources.
The R&D office of a pharma and chemical company. Three windows are open on the researcher's monitor: the internal technical-report system, a patent search tool, an external academic-paper database. To make a single judgment call, he toggles between these three all day long. Find a lead in a report, move to patents to check the scope of rights, then cross over to papers to align it with the latest research trends.
Finding the answer isn't the end. To report or cite which report and which page, which patent and which claim, which paper — he has to organize the sources all over again from scratch. Finding the answer and re-tracing the basis of that answer split into two rounds of labor.
What follows is the story of a customer PoC (proof-of-concept) and demo currently underway with one R&D organization. Query three lines of evidence from one channel, and bring the basis along every time an answer comes out.
Three kinds of evidence, scattered across three systems
In pharma and chemical R&D, a single judgment rarely stands on material from one place alone. Internal technical reports accumulated over time, patents that draw the lines of rights, the outside world's latest academic papers. Usually you have to look at all three at once. The problem is that these three mostly live in different systems and repositories.
- Internal technical reports. They pile up in the report system, and you find past documents by rummaging through them by keyword
- Patents. You find them by setting conditions in the patent search tool. The claims draw the scope of rights
- External academic papers. They live in an academic DB and carry metadata like title, author, source, and publication info
- Each system's search method and screen is different. To gather the basis for one question you have to toggle between systems several times, and cross-checking the pieces is the researcher's own job
- Internal resources alone make it hard to keep up with the fast-moving flow of the latest research. You want to see inside and outside in one flow, but the tools run separately
Four searches, one manual stitch
The old approach runs in roughly four steps. You rummage through past documents by keyword in the internal report system. You set conditions in the patent search tool to find related patents. You search papers in an external academic DB. Finally, a person directly cross-checks and organizes the evidence gathered from the three places.
This flow has two fractures. One is the keyword trap. If the search term diverges even slightly from the field's domain terminology, the very evidence that matters drops out of the results. Even for the same substance or the same process, if a report and a paper use different expressions, keyword search can't fill the gap between them. The other is the separation of sources. Since you have to note the answer and the source separately, when you later write a report or cite a conclusion, you wander around looking for where that was again.
I found the answer, but then I spend more time finding where that answer came from.
Put all three lines into one workspace
The setup is getting all three kinds of material into one workspace. Reports and patent documents go in as files, external papers as links. Your originals stay as they are, because the app reads a copy.
The order for loading reports, patents, and papers
Open "Extract sources"
Sidebar → "Extract sources" (or ⌘⇧I). Drag in the folders of internal technical reports and patent documents.
Paste papers through "Add URL"
Paste external academic-paper links through "Add manually" → "Add URL", one per line. Internal material and outside material land in the same list.
Press "Import"
The five-stage checklist starts. Before extraction, the "Getting to know your topics" stage pulls a domain vocabulary out of your own documents, so extraction fits your field's terminology. This is where the fracture — keywords diverging from domain terms and dropping evidence — gets filled.
You can ask mid-build
Send a chat message while extraction is running and the "Knowledge graph still building" dialog pops up. Pick "Send when extraction finishes" and your question goes out once the graph is fully stacked. Your choice is remembered for the rest of the build.
Move on to "Ask a question about your documents"
The "Your documents are ready" card offers two exits. Take "Ask a question about your documents" and the one channel for querying all three lines opens.
Ask all three lines from one channel
"What do the internal reports have on this substance?"
→ The agent searches the internal technical reports and answers.
The answer cites "which document, which page" alongside.
Conclusion and source come out together.
"Which patents catch this composition range and these process conditions?"
→ Not by simple keywords, but by real research conditions like
composition ranges and process conditions, it filters relevant patents.
It becomes the starting point for reviewing the scope of rights.
"How far has the outside research on this topic come?"
→ It researches papers based on title, author, source, and
publication-info metadata. The internal basis and the outside
world's latest research connect within one flow.
(after the three lines are gathered)
→ The gathered evidence and sources come back organized as a report.
You start the review from already-organized results, not a blank screen.Sources ride along with the answers because the graph has already read the documents. On an entity page, hover a connection chip and the source sentence behind the fact appears, cited as "{note} · line {n}". With conclusion and source arriving together, there's less need to note the source separately.
It doesn't throw all three lines at once with no order. It sets the priority of what to look at first, and handles repeated lookups fast by storing a result once it's fetched. It doesn't entrust everything to a single model either. Multiple LLMs (large language models) are being comparatively evaluated to home in on the configuration that best fits this work.
What changes
| Before | Now |
|---|---|
| Toggling all day between three windows: the report system, the patent tool, the academic DB | You ask reports, patents, and papers from one channel |
| If the search term diverges from domain terminology, key evidence drops out of the results | Extraction runs on the domain vocabulary that "Getting to know your topics" pulled out, and the goal is to fill that fracture with condition-based search and metadata research |
| You note the answer and the source separately, and hunt for where it was when you cite it | Every answer brings a document name, a page, or metadata along as its source |
| A person directly cross-checks and organizes the evidence gathered from three places | The gathered evidence and sources come back organized as a report |
What the PoC sets out to validate is this structure: gather three scattered kinds of evidence into one flow, and accompany every answer with its source. How much that structure raises the speed and trust of actual research judgment hasn't been measured yet — that's the stage the PoC is validating.