Having problems including linked PDFs in a site's inventory? If your role is either administrator or manager, you will be able to access the following settings.
Review the items listed below to help find the issue:
Is PDF checking enabled on the site? This option is found on the Advanced tab of the site's setting. If you don't see the option to Enable PDF checking, contact us so we can enable that feature for your account.
Are PDFs on your site located at a different base URL from what has been configured as the Start URL on the site? For example, consider a site configured with a Start URL of
https://admissions.college.edu, but the site's PDF files are stored athttps://resources.college.edu. In this case, you will need to add thatresourcesdomain as an Additional Domain Allowed in the Crawl in the Advanced tab of the site's settings for the PDFs to be included.
Do the PDF URLs have queries appended to them? For example:
https://resources.college.edu/pdfs/MeetingMinutes.pdf?v=20250104. If so, you need to ensure that Ignore pages with URL queries is not checked in the Ignored Paths tab of the site settings.
If the site is configured as a Site Map crawl, the PDF URLs must be included in the site map.
If the site has a Page Limit set, consider increasing that limit if the PDFs are found in the site's Crawl Log, but not included in the inventory.
If you have questions, please contact our DubBot Support team via email at help@dubbot.com or via the blue chat bubble in the lower right corner of your screen. We are here to help!

