html-extraction

Here is 1 public repository matching this topic...

RayenMalouche / MCP-PDF-Extractor-server

A Java-based server leveraging Apache Tika to extract content and metadata from files (PDF, DOCX, TXT, etc.) in a local files-to-extract directory. Supports HTML (with CSS styling) and text extraction, file listing, and metadata retrieval via MCP-compliant tools and REST APIs. Built with Spring Boot, Jetty, and MCP SDK.

java html pdf parser mcp extractor pdf-extractor html-extraction html-extractor pdf-extraction mcp-server modelcontextprotocol extractor-to-html

Updated Aug 30, 2025
Java

Improve this page

Add a description, image, and links to the html-extraction topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the html-extraction topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

html-extraction

Here is 1 public repository matching this topic...

RayenMalouche / MCP-PDF-Extractor-server

Improve this page

Add this topic to your repo