If this project saved you some time or made your day a little easier, a star would mean a lot — it helps others find it too.
The official Peppol Directory (PD; https://directory.peppol.eu).
This project is part of my Peppol solution stack. See https://github.com/phax/peppol for other components and libraries in that area.
This project is split into the following sub-projects:
phoss-directory-indexer- the PD indexer part (requires Java 25 since v0.17.0)phoss-directory-publisher- the PD publisher web application (requires Java 25 since v0.17.0)phoss-directory-client- a client library to be added to SMP servers to force indexing in the PD (requires Java 17)phoss-directory-searchapi- a client library for easier use of the Directory search REST API (since v0.7.2; requires Java 17)
Previous modules:
-
phoss-directory-businesscard- the common Business Card API - until v0.12.3; then moved to com.helger.peppol:peppol-directory-businesscard in https://github.com/phax/peppol-commons -
Production version is available at https://directory.peppol.eu (for Peppol)
- It can only handle participants registered at the SML
- For the indexing REST API, a client certificate (SMP production) is needed
-
Test version is available at https://test-directory.peppol.eu
- It can only handle participants registered at the SMK
- For the indexing REST API, a client certificate (SMP test) is needed
To build the PD software you need at least Java 25 and Apache Maven 3.x.
The two artifacts that are consumed by third parties - phoss-directory-client and phoss-directory-searchapi -
are compiled for Java 17, so that SMP servers running on Java 17 can keep using them.
All other modules only ever run inside the Directory server itself and are compiled for Java 25.
Additionally to the contained projects you MAY need the latest SNAPSHOT of ph-oton as part of your build environment.
The PD client is a small Java library that uses Apache HttpClient to connect to an arbitrary phoss Directory Indexer to perform all the allowed operations (get, create/update, delete).
The PD client uses ph-config to resolve configuration items.
See https://github.com/phax/ph-commons/wiki/ph-config for the details on the resolution logic.
Note: the old file pd-client.properties is not evaluated anymore.
Note: the configuration properties were heavily renamed in v0.10.0. Previous old names are shown in brackets.
The following configuration items are supported by the PD Client:
pdclient.keystore.type(old:keystore.type) (since v0.6.0) - the type of the keystore. Can beJKSorPKCS12(case insensitive). Defaults toJKS.pdclient.keystore.path(old:keystore.path) - the path to the keystore where the SMP certificate is containedpdclient.keystore.password(old:keystore.password) - the password to open the key storepdclient.keystore.key.alias(old:keystore.key.alias) - the alias in the key store that denotes the SMP keypdclient.keystore.key.password(old:keystore.key.password) - the password to open the key in the key storepdclient.truststore.type(old:truststore.type) (since v0.6.0) - the type of the keystore. Can beJKSorPKCS12(case insensitive). Defaults toJKS.pdclient.truststore.path(old:truststore.path) (since v0.5.1) - the path to the trust store, where the public certificates of the phoss Directory servers are contained. Defaults totruststore/pd-client.truststore.jkspdclient.truststore.password(old:truststore.password) (since v0.5.1) - the password to open the truststore store. Defaults topeppolhttp.proxy.host(old:http.proxyHost) - the HTTP proxy host for HTTP connections only. No default.http.proxy.port(old:http.proxyPort) - the HTTP proxy port forhttpconnections only. No default.- Removed in 0.10.0:
https.proxyHost- the HTTP proxy host forhttpsconnections only. No default. - Removed in 0.10.0:
https.proxyPort- the HTTP proxy port forhttpsconnections only. No default. http.proxy.username(old:proxy.username) (since v0.6.0) - the proxy username if http or https proxy is enabled. No default.http.proxy.password(old:proxy.password) (since v0.6.0) - the proxy password if http or https proxy is enabled. No default.http.connect.timeout.ms(old:connect.timeout.ms) (since v0.6.0) - the connection timeout in milliseconds to connect to the server. The default value is5000(5 seconds). A value of0means indefinite. A value of-1means using the system default.http.response.timeout.ms(old:http.request.timeout.msorrequest.timeout.ms) (since v0.10.3) - the response/request/read timeout in milliseconds to read from the server. The default value is10000(10 seconds). A value of0means indefinite. A value of-1means using the system default.https.hostname-verification.disabled(since v0.5.1) - a boolean value to indicate if https hostname verification should be disabled (true) or enabled (false). The default value istrue.
Example PD Client configuration properties:
# Key store with SMP key (required)
pdclient.keystore.type = pkcs12
pdclient.keystore.path = smp-test.p12
pdclient.keystore.password = password
pdclient.keystore.key.alias = cert
pdclient.keystore.key.password = password
# Default trust store (optional)
pdclient.truststore.type = pkcs12
# For Test:
pdclient.truststore.path = truststore/2025/smp-test-truststore.p12
# For production:
# pdclient.truststore.path = truststore/2025/smp-prod-truststore.p12
pdclient.truststore.password = peppol
# TLS settings
https.hostname-verification.disabled = falseThe PD Indexer is a REST component that is responsible for taking indexing requests from SMPs and processes them in a queue
(Peppol SMP client certificate required).
Only the Peppol participant identifiers are taken and the PD Indexer is responsible for querying the respective SMP data directly.
Therefore the respective SMP must have the appropriate Extension element of the service group filled with the business
information metadata as required by PD.
Please see the PD specification
for a detailed description of the required data format as well as for the REST interface.
Configuration properties:
indexer.maxparallel- The number of work items that are indexed in parallel, and therefore the number of SMP queries that are performed in parallel. Defaults to4.
Handling a single work item is dominated by waiting for the SMP to respond and not by CPU usage, and since v0.17.0 the indexing threads are virtual threads. This value may therefore be raised way beyond the number of available cores. Be aware that raising it directly increases the request rate that the Directory puts on the SMPs of the network.
The PD Indexer supports shadowing of indexing requests to a downstream replicator service for migration purposes (e.g., PD2 migration). When enabled, successful indexing requests are replicated asynchronously as custom JSON events to a configured downstream URL.
Configuration properties:
indexer.shadowing.enabled- Enable or disable indexer request shadowing. Defaults tofalse.indexer.shadowing.url- The downstream URL to send shadow events to. Required if shadowing is enabled.indexer.shadowing.timeout.ms- HTTP timeout in milliseconds for shadow requests. Defaults to5000(5 seconds).indexer.shadowing.interval- Interval in which the dispatcher job processes the queued events. Defaults to1m(1 minute). The value uses the duration grammar (e.g.30s,5m,1h 30m). The deprecated propertyindexer.shadowing.interval.secondsis still evaluated as a fallback.indexer.shadowing.checkpoint- Interval after which the complete shadow event list is written to disk. Defaults to5m(5 minutes). The value uses the duration grammar (e.g.30s,5m,1h 30m). The write-ahead log provides durability in between, so a larger value only results in a longer WAL replay on startup.indexer.shadowing.secret- Optional secret string included in theX-Shadow-SecretHTTP header for authentication. If not set, no authentication header is sent.
Shadow event format:
Shadow events are sent as HTTP POST requests with JSON payload containing:
eventId- Unique UUID for idempotencycreatedAt- ISO 8601 timestampoperation- Operation type (CREATE_UPDATEorDELETE)participantId- The participant identifierrequestingHost- The requesting hostclientCertificate- Object containing:sha256Fingerprint- SHA-256 fingerprint of the client certificate (primary identity)subjectDN- Certificate subject distinguished nameissuerDN- Certificate issuer distinguished name
Operational notes:
- Shadow events are persisted to disk (
shadow-events.xml) before dispatch for crash safety - A background job dispatches events at the configured interval (default: every minute)
- Failed events with non-retryable errors (HTTP 4xx) are moved to a dead-letter queue (
failed-shadow-events.xml) - Failed events with retryable errors (network issues, HTTP 5xx) remain in the queue for automatic retry
- Shadow failures never affect the original indexing request
- Assumes single application instance per data directory
- Optional authentication via
X-Shadow-Secretheader for securing the downstream endpoint
Example configuration:
# Indexer shadowing (disabled by default)
indexer.shadowing.enabled=false
indexer.shadowing.url=https://pd2-replicator.example.com/shadow-events
indexer.shadowing.timeout.ms=5000
indexer.shadowing.interval=1m
indexer.shadowing.checkpoint=5m
indexer.shadowing.secret=your-secret-token-hereThe PD Publisher is the publicly accessible web site with listing and search functionality for certain participants.
v0.18.2 - 2026-09-09
- The country selector of the search page offers all countries known to the Java runtime, sorted alphabetically by their display name, instead of only the countries of the country specific Peppol participant identifier schemes. The country that is searched for is the country of a Business Card and is therefore not limited to those schemes
- The class
HCPeppolCountrySelectwas renamed toHCCountrySelectand its methodgetAllPeppolCountries ()togetAllCountries () - The list is based on
Locale.getISOCountries (). The country cache of ph-commons is not used, because it is filled from the available locales and therefore also contains the UN M.49 region codes like419(Latin America) that are no countries
- The class
- The
GETrequests of the/exportAPIs are answered with HTTP 406 (Not Acceptable) if the client does not accept the GZIP content encoding. The export files are huge, so they are only handed out to clients that can take them compressed- A request is rejected if its
Accept-Encodingheader is present and does not accept GZIP - likeidentity,br,deflate,gzip;q=0,*;q=0or an empty value.gzip,x-gzipand*with a quality above 0 are accepted - A request without an
Accept-Encodingheader states no preference at all, so according to RFC 9110 section 12.5.3 the choice of the content coding is up to the server. Such a request is therefore served and not rejected - The check happens before the rate limit is consumed, so a rejected request does not use up one of the few daily download slots
- The
HEADrequests, that only deliver the export metadata, are not affected, because their responses carry no content anyway
- A request is rejected if its
- The public documentation pages of the publisher now describe multilingual Business Entity names
- The JSON example of the "REST API documentation" page uses the real structure of the
namefield of a Business Entity - an array of objects with the mandatory fieldnameand the optional fieldlanguage- instead of the plain string that was shown before. One Business Entity of the example carries its name in two languages, the other ones show the far more common case of a single name without a language - The XML example of the same page shows the optional
languageattribute of thenameelement as well - The "Export data" page now contains an example of the Business Card JSON export, that includes a Business Entity with a multilingual name, and it documents the fields of the top-level object of that export
- The JSON example of the "REST API documentation" page uses the real structure of the
v0.18.1 - 2026-09-08
- Fixed several paths that prevented entries of the re-index list from ever being retried
- The retry period of a re-index work item is now anchored on the moment the item enters the re-index list, instead of on the creation date time of the underlying indexer work item. The time an item spent in the indexer work queue - which for a bulk indexing or across a server downtime may exceed
reindex.maxretryhours- was previously deducted from the retry period, so that such items were moved to the dead list before their first retry was even due PDIndexerManager.reIndexParticipantDataSynchronously ()takes all due items off the re-index list before it handles them one by one. An unexpected error in a single item aborted the loop, so that all the remaining items of that run were silently lost - neither in the re-index list nor in the dead list, but still in the internal "unique items" list, which blocked the affected participants from ever being indexed again. Each item is now handled separately and is put back into the re-index list if it could not be handled. The same applies toPDIndexerManager.expireOldEntries ()and the dead listReIndexWorkItemList.getAndRemoveAllEntries (...)andReIndexWorkItemList.getAndRemoveEntry (...)searched the items outside of the write lock that removed them. A concurrent removal - e.g. from a newly queued work item or from deleting participants - therefore led to an item being handed out to two callers at once. The search and the removal now happen within a single write lock, and only items that were really removed are returned- The work items of the re-index list are restored as "in progress" upon startup. The respective list is in memory only, so a re-index entry did not prevent a new work item for the same participant from being queued in parallel after a restart, and the re-index list could end up with several entries for the same participant
- The retry period of a re-index work item is now anchored on the moment the item enters the re-index list, instead of on the creation date time of the underlying indexer work item. The time an item spent in the indexer work queue - which for a bulk indexing or across a server downtime may exceed
- The re-index job is scheduled as the last action of
PDIndexerManager.startIndexing (), so that it cannot start working on a partially restored state
v0.18.0 - 2026-09-07
- The public documentation pages of the publisher were updated to the current state of the implementation
- The "Export data" page documents the participant identifier exports (
/export/participants-xml,/export/participants-jsonand/export/participants-csv), the per IP and per file rate limiting of the downloads and that the download URLs respond with an HTTP redirect to the storage location - The "How to use it" page no longer declares the REST API and the data download as "work in progress"
- The "REST API documentation" page describes that an invalid or too short search term is ignored instead of leading to an HTTP status code 400, that every request parameter may occur more than once, and uses current Peppol document type identifiers in all examples
- The "Export data" page documents the participant identifier exports (
- The "export all" job now really starts at 02:00 a.m. UTC, as it is documented, instead of at 02:00 a.m. in the time zone of the server
- A search no longer queries the search index twice.
IPDIndex.searchAll (...)returns the total number of matching documents, that every search engine determines as a side effect of the search itself, instead ofvoid. The separategetCount (...)call that the search UI and the REST search API used to fill in the total result count is therefore gonePDStorageManager.searchAll (...)andPDStorageManager.searchAllDocuments (...)return the total hit count as well- The new method
PDStorageManager.search (...)returns the new recordPDSearchResult, that contains the matching business entities as well as the total hit count.PDStorageManager.getAllDocuments (...)is unchanged and is now a shortcut for it - Because both numbers now originate from the same query, the displayed total result count can no longer belong to a different state of the index than the returned entries
- Incompatible change: implementations of
IPDIndexmust change the return type ofsearchAll (...). Callers that are not interested in the total hit count can ignore the return value and don't need to be changed
- The OpenSearch implementation enables the exact total hit count tracking for searches with a maximum result count, because OpenSearch only counts up to
index.max_result_window(10.000 by default) hits otherwise - The export endpoints answer
HTTP HEADrequests with the metadata of the currently published export file, so that a consumer can find out whether a new export exists, without spending one of its few daily download slots- The response uses HTTP status code 200 and contains the
ETagand theLast-Modifiedof the stored file, its size in the new headerX-Export-Content-Lengthand its content type in the new headerX-Export-Content-Type. It deliberately does not issue the redirect thatGETissues, because a signed redirect would be equally usable for a download and would therefore create an unmetered download path - See the new class
ExportMetadataHttpHandlerthat replaces the defaultHEADhandler ofExportServlet - Because a
HEADrequest transfers no content, it has its own, far more generous rate limit per IP and per file, that is configurable via the new propertyexport.limit.headrequestsperday(default 100) - see the new methodExportRateLimit.isOverHeadLimit (...) - The metadata read from S3 is cached in memory for the duration of the new property
export.metadata.cache(default 5 minutes, duration grammar), so that only the firstHEADrequest per file and interval causes an S3 round trip - see the new methodS3Helper.headS3Object (...)
- The response uses HTTP status code 200 and contains the
- The
Cache-Controlheader of the export redirects isno-storeinstead ofmax-age=86400, if the signing of the export URLs is enabled. The redirect then carries a short lived signature, so caching it would hand out a URL that is rejected as soon as the signature expired - The "Export data" documentation page describes the
HEADrequests, and that the redirect itself must not be cached and must not be reused, whereas the downloaded data may still be cached for up to 24 hours - A search for a country code now also finds the entries that use a synonymous country code, because a Business Card may use a country code that differs from the ISO 3166-1 alpha-2 code of the respective code list.
See issue #15 - thx @clancger
- "GB" and "UK" are treated as synonyms of each other, as are "GR" and "EL" (the EU VAT prefix of Greece). Searching for any of them returns the entries of all the codes of the group
PDQueryManager.getCountryCodeQuery (...)therefore creates a query that matches any of the synonymous country codes, instead of a single exact match. It is unchanged for every country code that has no synonyms- This applies to the REST API search parameter
countryas well as to the country selector of the search page
- The background image of the public search page is a JPEG instead of a PNG -
imgs/peppol/peppol.jpg(66 KB) replacesimgs/peppol/peppol.png(941 KB), reducing the largest asset of the search page by 93%- The image is a photo, for which the PNG format is inappropriate. It is stored as a progressive JPEG without metadata, so a browser can already display it while it is still being transferred
- The dimensions are unchanged (2180x520 pixels), so the image stays sharp on high resolution displays. The CSS class
big-query-imagefills the Bootstrap container, that is at most 1320 pixels wide, usingbackground-size: cover
v0.17.3 - 2026-09-05
- The interval after which the shadow event list is written to disk in total is configurable via the new property
indexer.shadowing.checkpoint. It defaults to 5 minutes and uses the duration grammar (e.g.30s,5m,1h 30m), soPDServerConfiguration.getIndexerShadowingCheckpointDuration ()returns aDuration- Previously each checkpoint rewrote the complete file every 10 seconds, which is heavily disproportionate to the amount of data that actually changed if the queue is large. The write-ahead log keeps the events durable in between
- The interval of the shadow event dispatcher uses the duration grammar as well and is therefore configured via the new property
indexer.shadowing.interval(e.g.1m) instead ofindexer.shadowing.interval.seconds. The new methodPDServerConfiguration.getIndexerShadowingIntervalDuration ()returns aDurationand replacesgetIndexerShadowingIntervalSeconds ()- The deprecated property
indexer.shadowing.interval.secondsis still evaluated, if the new one is not present, and logs a warning once. It will be removed in a future version
- The deprecated property
- The search results of the name search and of the generic search are now ordered by relevance (see issue #49)
- An entry in which a word is exactly the search term is ranked highest, followed by the entries in which a word starts with the search term, followed by the entries that merely contain the search term inside a word
PDQueryManager.getNameQuery (...)therefore uses the previous "contains" query as a mandatory clause that does not contribute to the score, and adds a term query and a prefix query per search term as optional clauses that only influence the ordering- The new method
PDQueryManager.getGenericQuery (...)does the same for the generic search fieldq. Additionally a match in the name of a business entity is ranked higher than a match in any other field - The search fields that perform an exact match (participant ID, country, identifier scheme, identifier value, registration date and document type ID) are combined as filters, so that they only limit the result set but no longer influence the ordering. See the new method
EPDSearchField.getCombinationOccurrence (). If a query consists of such search fields only, all clauses are combined as before, so that the ordering of these queries is unchanged - The set of matching entries is not affected by any of these changes - a mandatory clause selects the same documents no matter if it contributes to the score, and the added optional clauses can never add or remove a match. Both properties are asserted by the conformance test suite
IPDIndex.searchAll (...)must now return the documents ordered by descending relevance, if a positive maximum result count is provided. Both the Apache Lucene and the AWS OpenSearch implementation already did that
- The DataTables of the pages "Index Queue", "Re-Index List" and "Dead Index List" use server side pagination, so that only the rows of the currently displayed page are rendered
- Previously every one of these tables was rendered as a whole and the result was kept in the session, which means the memory consumption was proportional to the number of entries times the number of logged in users
- Paging, sorting and searching are performed on the work items themselves. The new enums
EIndexerWorkItemColumnandEReIndexWorkItemColumntie each shown column to the respective comparator and to the value the global search is performed on - The date and the number columns are sortable but no longer searchable, because the global search would have to match the localized text shown in the respective cell
- The search page offers a country selector next to the query field, that limits the results to a single country. The default selection is "All countries", meaning that no country filter is applied at all
- The selectable countries are the countries of all the country specific Peppol participant identifier schemes - see the new class
HCPeppolCountrySelectthat is based onPeppolParticipantCountryHelper.getAllSchemeCountryCodes () - The selected country is passed in the request parameter
country- the same name the REST API search uses - and is applied on the country code of the business entity. Being an exact match, it is combined as a filter and therefore does not influence the relevance ordering of the results
- The selectable countries are the countries of all the country specific Peppol participant identifier schemes - see the new class
- The job that exports all data creates audit items for its start, for both of its intermediate steps and for its end - the audit actions are
export-all-start,export-all-participant-ids,export-all-formatsandexport-all-end(see the constants inExportAllDataJob)- Every one of these audit items carries the overall duration of the job up to that point, so that the audit trail alone shows how long the export took. The two intermediate items carry the duration of their respective step as well
- A step that failed is audited as a failure, including the names of the export formats that could not be created or uploaded
- The bulk import and the bulk delete of participants create an audit item when they start and when they end -
index-import-start/index-import-endandindex-delete-start/index-delete-end, built from the job type via the new methodAbstractPDParticipantFileJob.getAuditAction (...)- Both items carry the ID of the user that triggered the job - the job runs in a worker thread, so the audit item itself is not bound to that user - plus the name of the uploaded file and the overall duration
- The "end" item is audited as a failure if the job failed. If the job could not even be started, the "start" item is audited as a failure by the upload page
AbstractPDParticipantFileJob.createLongRunningJobResult ()is final now and merely wraps the new abstract methodcreateParticipantJobResult (), so that the outcome of the job is known when the "end" audit item is created
- The DataTables of the page "Participant list" uses server side pagination as well - exactly like the pages "Index Queue", "Re-Index List" and "Dead Index List" do
- The new class
ContainedParticipantis the row type, and the new enumEContainedParticipantColumnties each shown column to the respective comparator and to the value the global search is performed on. The participant ID is sortable and searchable, the entity count is sortable only - The parameters
showallandmaxentriesas well as the limit of 500 rendered participants were removed - every participant is reachable via paging now - The heading with the number of participants was removed, because the DataTables shows the total number of entries anyway, and determining it needs a full query of the search index
- Note that every paging, sorting and searching operation queries all participants from the search index anew - the same query the previous rendering of the whole page performed once
- The new class
v0.17.2 - 2026-09-04
- Added the page "Bulk delete participants" to delete participants from the search index from an uploaded file
- The entries are deleted without verifying the owner, exactly like on the "Manually delete participant" page, so that entries owned by an SMP can be removed as well
- All pending indexer work items of the deleted participants are withdrawn as well, because an already queued create/update work item would otherwise put the participant right back into the index
- Added
PDIndexerManager.removeWorkItems (...)that removes all work items of a set of participants from the work queue, the re-index list and the dead list
- Renamed the page "Import participants" to "Bulk import participants". Its menu item ID is unchanged, so existing bookmarks keep working
- Both bulk pages show a toolbar that links to the "Long running jobs" page, if that page is visible to the logged in user
- The "Bulk import participants" page no longer performs the import in the HTTP thread, so that importing tens of thousands of participants does not block the browser anymore
- The uploaded file is stored below the data path and the new long running job
PDIndexImportJobreads it and queues the participants in the background. The outcome is shown on the "Long running jobs" page - Only a single import may run at a time, and the page indicates whether one is currently running
- The result no longer lists every single participant, but the number of queued, duplicate, already queued and syntactically invalid participant IDs, plus the first 100 entries of the problematic ones
- Duplicate participant IDs within the import file are now detected and reported separately from those that already are in the indexing queue
- The uploaded file is stored below the data path and the new long running job
- Both participant file pages accept the same two file formats, and the format is detected from the content of the file
- XML - every element called
participantwith the attributesschemeandvalue, no matter where in the document it appears. That covers the participant list export as well as the full Business Card export - Text - one URI encoded participant ID per line (e.g.
iso6523-actorid-upis::9915:test). Lines are trimmed, empty lines and lines starting with#are ignored
- XML - every element called
- Added
PDIndexerManager.queueWorkItems (...)to queue many participants at once. Queueing n participants scanned the re-index list and the dead list n times each and wrote one log line per participant, which made bulk imports unusable. Both lists are now scanned exactly once per bulk call - Added the button "Re-index all entries now" to the "Dead Index List" page, to move all dead entries back into the indexing queue. See #89
- Each entry is queued with its original action type, so that the respective entry is also removed from the dead list
- Fixed that on startup the persisted indexer work items were read and executed before the Business Card provider was set, so that all of them failed with "No BusinessCard Provider is present." and were moved to the re-index list. See #90
- The
PDIndexerManagerconstructor no longer starts any indexing activity - the new methodPDIndexerManager.startIndexing ()schedules the re-index job and reads the persisted work items, and it must be called after the Business Card provider was set startIndexing ()throws anIllegalStateExceptionif no Business Card provider is present, so that a wrong startup order cannot happen unnoticed
- The
- The export of all Business Cards and participants queries the search index only once per participant instead of once per participant and export format, reducing the number of index queries of a full export run by 75%. See #88
- All the export formats are now created in a single pass over the participants, so that each participant is read, parsed and converted to
PDStoredBusinessEntityobjects only once - Backwards incompatible change: the
ExportAllManager.writeFile*methods were replaced byExportAllManager.exportAll (...)that takes the export formats to be created - Added the interface
IExportAllHandlerand the base classAbstractExportAllHandler- one implementation per export format, all fed with the data of every participant - Fixed that an export format that failed to be created was nevertheless uploaded to S3, overwriting the previously good file with a truncated one. Now only successfully created files are uploaded
- The export status shown in the administration UI now contains the export progress and the name of the export format that is currently uploaded
- All the export formats are now created in a single pass over the participants, so that each participant is read, parsed and converted to
v0.17.1 - 2026-08-31
- Fixed a concurrency issue in the Apache Lucene search index that could return the Business Card of an unrelated participant, if the index was modified while a search was running (security advisory GHSA-8qhv-6p5x-2437)
- The internal Lucene document IDs of a search result are only valid for the index reader that created them, but they were resolved via an independently obtained index reader that may have been reopened in the meantime
PDLucenenow uses a LuceneSearcherManager, so that all the documents of a search result are resolved with exactly the searcher that was used for searching, and so that an index reader that is in use is neither replaced nor closed underneath the caller- This also fixes that the previously used
DirectoryReaderwas never closed when the index changed - Backwards incompatible change:
PDLucene.getDirectoryReader (),PDLucene.getSearcher ()andPDLucene.getDocument (int)were replaced byPDLucene.acquireSearcher ()andPDLucene.releaseSearcher (IndexSearcher)- the previous API could not be used in a thread-safe way - Backwards incompatible change: the interface
ILuceneDocumentProviderwas removed and the constructor ofAllDocumentsCollectortakes the document consumer only - the documents are now resolved from the leaf reader that is currently being searched - The
phoss-directory-indexer-opensearchimplementation was not affected - Added the test
PDLuceneIndexConcurrentSearchFuncTestthat searches in parallel to index modifications and verifies that every result belongs to the queried participant
v0.17.0 - 2026-08-30
- The modules that only ever run inside the Directory server itself are now compiled for Java 25:
phoss-directory-indexer,phoss-directory-indexer-lucene,phoss-directory-indexer-opensearch,phoss-directory-indexer-conformanceandphoss-directory-publisherphoss-directory-clientandphoss-directory-searchapiare still compiled for Java 17, so that SMP servers running on Java 17 can keep using them- Building the project therefore requires Java 25 - the new POM property
java.version.serverholds the version of the server side modules
- Backwards incompatible change:
IPDIndexQueryis now asealedinterface that only permits the contained implementations, andPDIndexQueryBool,PDIndexQueryContains,PDIndexQueryPrefixandPDIndexQueryTermare nowfinal. The set of queries was documented as being closed since v0.16.0 - it is now enforced by the compiler- As a result
PDLuceneIndexandPDOpenSearchIndextranslate the queries with an exhaustiveswitchinstead of a chain ofinstanceofchecks, so that adding a new query type is a compile error in every implementation ofIPDIndexinstead of a runtime error in the one that was forgotten - The same applies to
EPDIndexQueryOccur- adding a new occurrence is now a compile error in both implementations
- As a result
- The indexer work items are now handled by virtual threads instead of a fixed pool of 4 platform threads
- Added the new configuration property
indexer.maxparallelto configure the number of work items that are handled in parallel. It defaults to4, so the load that the Directory puts on the SMPs of the network is unchanged unless the property is raised deliberately - Backwards incompatible change: the constructor of
IndexerWorkItemQueuetakes the number of parallel work items as the second parameter - Added
IndexerWorkItemQueue.getMaxParallel () - Added
PDServerConfiguration.getIndexerMaxParallel ()and the constantPDServerConfiguration.DEFAULT_INDEXER_MAX_PARALLEL - Note: virtual threads are always daemon threads, whereas the previous indexer threads were not.
PDIndexerManager.close ()stops the queue explicitly on shutdown, so the remaining work items are still persisted - The administration navbar shows the configured parallelism next to the index queue length
- Added the new configuration property
PDIndexExecutortranslatesEIndexerWorkItemTypewith an exhaustiveswitchexpression, so that adding a new work item type is a compile error- Fixed that "Download results as XML" on the public search page ignored the maximum result count of the search (default 50, see the
maxquery parameter) and exported all matching Business Cards instead- Backwards incompatible change:
ExportAllManager.queryAllContainedBusinessCardsAsXMLtakes the maximum result count as the second parameter - Backwards incompatible change:
PDSessionSingleton.setLastQuerytakes the maximum result count as the second parameter - Added
PDSessionSingleton.getLastQueryMaxResultCount ()
- Backwards incompatible change:
- Note: when running on Java 25, Apache Lucene 8.11.4 logs a warning that
sun.misc.Unsafe::invokeCleaneris terminally deprecated, because that is howMMapDirectoryunmaps index files. This is harmless today but the method will be removed in a future JRE - the OpenSearch based search index is not affected
v0.16.0 - 2026-08-19
- The publisher web UI was switched from Bootstrap 4 to Bootstrap 5, using
ph-oton-bootstrap5instead ofph-oton-bootstrap4- No functional change - this is a pure UI framework update
- All search index access is now performed via the new search engine independent interface
IPDIndex(packagecom.helger.pd.indexer.searchindex), so that a different search engine can be plugged in later- Added the search engine independent document model
PDIndexDocument/PDIndexFieldreplacing the usage of the LuceneDocumentandFieldclasses - Added the search engine independent query model
IPDIndexQuery(packagecom.helger.pd.indexer.searchindex.query) replacing the usage of the LuceneQueryandTermclasses - Added
PDLuceneIndexas the single implementation ofIPDIndex, that continues to use Apache Lucene as before - all Lucene usage is now limited to the packagecom.helger.pd.indexer.lucene - Each
IPDIndexQuerycaches the search engine specific query created from it, so that executing the same query object twice (as the UI and the REST API do, to get the results and the total hit count) does not translate it twice - creating the LuceneWildcardQueryobjects of a generic search costs roughly 0.5 milliseconds - The
PDQueryManagermethodsgetXYZLuceneQuerywere renamed togetXYZQueryand returnIPDIndexQueryinstead of the LuceneQuery - The
PDStringFieldmethodsgetExactMatchTermandgetContainsTermwere replaced bygetExactMatchQuery,getPrefixQueryandgetContainsQuery PDMetaManager.getLucene ()was replaced byPDMetaManager.getIndex ()PDStorageManager.searchAtomic (Query, Collector)was removed, because it was Lucene specific - usesearchAllinstead- No functional change - the index content and all search results stay the same
- Added the search engine independent document model
- The search index implementation is now pluggable and is selected via the new configuration property
searchindex.type- see docs/opensearch.md- Each implementation registers itself via the new SPI interface
IPDIndexProviderSPIand is resolved by the new classPDIndexFactory - Backwards incompatible change: the Apache Lucene implementation was moved from
phoss-directory-indexerto the new submodulephoss-directory-indexer-lucene(searchindex.type=lucene, still the default). The modulephoss-directory-indexercontains no search index implementation anymore, so exactly one implementation submodule must be added to the classpath.phoss-directory-publisherdepends onphoss-directory-indexer-lucene, so the default deployment is unchanged
- Each implementation registers itself via the new SPI interface
- Added the new submodule
phoss-directory-indexer-opensearchcontainingPDOpenSearchIndex, an implementation ofIPDIndexfor AWS OpenSearch (searchindex.type=opensearch)- The OpenSearch endpoint is configured via the new properties starting with
opensearch. - Supported authentication types are AWS Signature Version 4 (for managed AWS OpenSearch Service domains) and no authentication (for local testing) - HTTP basic authentication is not supported
- Unlike Apache Lucene, OpenSearch cannot delete and add documents atomically, and it does not guarantee the order in which the business entities of a participant are returned
- The OpenSearch endpoint is configured via the new properties starting with
- Added the new submodule
phoss-directory-indexer-conformancecontaining the search engine independent conformance test suite that every implementation ofIPDIndexmust passAbstractPDIndexConformanceTestasserts theIPDIndexcontract,AbstractPDStorageManagerConformanceTestassertsPDStorageManageron top of it- The previous
PDStorageManagerTestandPDLuceneIndexTestwere folded into these classes, so the very same assertions now run against Apache Lucene and AWS OpenSearch - The OpenSearch conformance tests skip themselves if no OpenSearch is reachable at
http://localhost:9200
v0.15.7 - 2026-08-05
- Updated to parent-pom 3.1.0, enabling Reproducible Builds
- Updated to ph-commons 12.3.3
- Updated to BouncyCastle 1.85, fixing a myriad of CVEs
- JSON writing now escapes all control characters (
U+0000-U+001F) as\u00XXper RFC 8259 - relevant for the JSON export - External XML resource resolution no longer resolves remote URL schemes by default, to prevent Server Side Request Forgery (SSRF)
- The internal soft maps used for caching are now thread-safe - previously concurrent reads could corrupt them and let cache eviction hang while holding the cache write lock
- Updated to ph-web 11.4.3
- Incoming requests are now wrapped in a
SafeHttpServletRequest, avoiding "The request object has been recycled" errors e.g. from the long running request monitor - The HTTP proxy configuration is now activated by the presence of
http.proxy.hostandhttp.proxy.portalone;http.proxy.enabledonly acts as an explicit kill-switch when set tofalse
- Incoming requests are now wrapped in a
- Updated to ph-oton 10.3.0 and ph-oton-bootstrap4 10.2.0
- Added throttling on login, if unknown user names are used
- Updated to Jetty 12.1.10
- Updated to peppol-commons 12.6.1
- Updated to Peppol eDEC Code Lists v9.7
- The SMP client now verifies that the participant and document type identifiers contained in the SMP response match the requested ones, so SMPs that resolve identifiers case insensitively are detected during indexing
- Removed the EC SML fallback in
EPeppolNetwork peppol-smp-clientno longer requires a JAX-WS (Metro) runtime
- Updated to peppol-ui 0.9.19
PDClientnow logs the absolute URL of each invoked indexer request
v0.15.6 - 2026-05-18
- Improved the verbosity and high-load handling of the internal scheduler - hoping we're capturing the underlying issue
v0.15.5 - 2026-05-18
- Added indexer shadowing support to replicate successful indexing requests to a downstream service (e.g. for PD2 migration). See the "Indexer Shadowing Configuration" section above for details
- Added
Cache-Control: max-age=86400header on all/export/*redirect responses to improve cacheability - Added per-IP per-file rate limiting on export redirects (sliding 24h window, configurable via
export.limit.requestsperday, default 3 requests) - Enabled S3 multipart uploads in
S3Helper - Removed OSGI bundling from
phoss-directory-client,phoss-directory-indexerandphoss-directory-searchapi
v0.15.4 - 2026-03-17
- REST API returns errors as valid
application/jsonorapplication/xmland no longer astext/plain - Added new configuration property
peppol.lookup.enabled - Added link to Peppol Lookup service if the search term is a participant ID (if enabled)
- Added link to lookup service also in the "Support" menu
v0.15.3 - 2026-02-13
- Showing the errors found during indexation on the UI for better support
v0.15.2 - 2026-01-23
- S3 buckets get a max-age of 6 hours to improve cachability
- The check for Peppol Participant Identifier Value syntax was improved to follow the rules from the Peppol Policy for use of Identifiers 4.4.0
v0.15.1 - 2026-01-19
- Requires at least Java 17 again - my bad
- Updated changelog
- Trying to get
Content-Dispositionto work
v0.15.0 - 2026-01-18
- Requires at least Java 21
- Instead of streaming the export files to local disk, they are now stream to S3
- Instead of reading the files from local disk, they are redirected to S3
v0.14.10 - 2025-12-30
- Updated to Peppol eDEC Code Lists v9.5
- Added new configuration propery
webapp.api.allow.originto configureAccess-Control-Allow-Originresponse header - Removed support for the
pd.propertiesandprivate-pd.propertiesconfiguration sources - The indexation happens now in 4 parallel threads
- The participant details are now showing the Participant Identifier Scheme name instead of the agency providing it
- The search result list details (like Name and Country) are now consistently aligned between the different participants
- Added a small hint for invalid Belgian CBE numbers in the Participant Details
- Fixed an error in the
directory-export-v3.xsd - Implemented an initial version of the change log. See #76
- Showing the current queue length in the title bar to more easily evaluate possible delays
v0.14.9 - 2025-11-16
- Updated to ph-commons 12.1.0
- Using JSpecify annotations
v0.14.8 - 2025-11-13
- Fixed an error, that document types were not correctly extracted, if the SMP response contains a percent encoded participant identifier (fixed in peppol-smp-client 12.1.2)
- Trying to disable DNSJava caches. See #77
v0.14.7 - 2025-11-04
- Created an updated XML Schema v3 for the XML export
- Using stream based JSON and XML export to reduce memory usage during export
v0.14.6 - 2025-11-03
- Improved "Export all" handling so that the participant list is queried only once
- Also removed any locking on export, to avoid blocking access to the export data while exporting
v0.14.5 - 2025-11-03
- Fixed a potential
NullPointerExceptionif a participant identifier could not be parsed - Internal ownership representation was changed to not use the serial number anymore, therefore deletion should also work after a certificate update
v0.14.4 - 2025-11-03
- Fixed resilience when loading stored values that are invalid identifiers - was blocking the export
v0.14.3 - 2025-11-02
- Updated to eDEC Code Lists v9.4
- In case of an HTTP 429 response, the
Retry-Afterheader is set to the seconds to wait - Removed unwanted "Peppol " in front of some predefined document type names
- Improved internal error and progress handling for export all job
- Made sure the HTTP 429 response is properly documented on the REST API documentation page
v0.14.2 - 2025-10-07
- Fixed HTTP response charset of CSV exports
v0.14.1 - 2025-10-03
- The export job is now scheduled to happen on 2am - more deterministically
- Added support for Peppol G2 + G3 support in parallel
- Removed the contact page form, as it was not working anymore
- Removed public page login
- Updated to eDEC Code Lists v9.3
- Fixed links to peppol.org and updated spelling where necessary
v0.14.0 - 2025-08-27
- Requires Java 17 as the minimum version
- Updated to ph-commons 12.0.0
- Removed all deprecated methods marked for removal
v0.13.6 - 2025-05-14
- Updated dependencies
- The version of the exported files have changed, because the code list states were incorporated (XML: v2 -> v3; JSON: v1 -> v2)
v0.13.5 - 2024-08-22
- Added a CORS HTTP response header for the REST API. See #68
v0.13.4 - 2024-07-30
- Updated to peppol-commons 9.5.0 with eDEC Code Lists v8.9
v0.13.3 - 2024-05-24
- Updated to peppol-commons 9.4.0
v0.13.2 - 2024-04-02
- Ensured Java 21 compatibility
v0.13.1 - 2024-03-22
- Fixed the
nameREST API query parameter
v0.13.0 - 2023-11-13
- Removed submodule
phoss-directory-businesscardand usingpeppol-directory-businesscardfrom https://github.com/phax/peppol-commons instead - Updated code lists to v8.7
v0.12.3 - 2023-10-27
- Fixed the name of the attribute for the client certificate retrieval (
jakarta.) - Added special handling for Peppol Wildcard identifiers on the UI
v0.12.2 - 2023-08-24
- Updated to ph-oton 9.2.0
- Updated code lists to v8.6
v0.12.1 - 2023-08-16
- Introducing class
PDResultListMarshallerin favour ofPDSearchAPI(Reader|Validator|Writer) - Added a BusinessCard JSON export
v0.12.0 - 2023-02-25
- Using Java 11 as the baseline
- Using Servlet API 5.0.0 as the baseline: JakartaEE 9, Java 11+, Apache Tomcat v10.0.x, Jetty 11.x
- Updated to Jersey 3.1.1
- Updated to ph-commons 11
- Updated the known names to eDEC Code List v8.3
v0.11.1 - 2025-10-01
- Add the Disclaimer on the website
- Update the known document type and process IDs to eDEC codelist v9.3
- Added rudimentary support for Wildcard identifiers
- Fixed a bug in the name search
- Removed the Twitter links
- Added a BusinessCard JSON export
- The export of data is constantly scheduled to 2am instead of startup time
- Added a CORS HTTP response header for the REST API. See #68
v0.11.0 - 2022-12-19
- Updated to Lucene 8.x
v0.10.5 - 2022-11-25
- Improved logging of indexation
- SMP client configuration became more resilient
v0.10.4 - 2022-11-14
- Added new configuration parameter
smp.tls.trust-allto disable the TLS certificate checks for the SMP client
v0.10.3 - 2022-08-17
- Updated to Apache Http Client v5.x
- Updated to ph-web 9.7.1
- Fixed an error in the REST API with the "name" parameter when multilingual names are used
v0.10.2 - 2022-03-28
- Removed the code for the handling of objects marked as deleted
- Improved the owner check upon deletion
v0.10.1 - 2022-03-09
- Added export of participant IDs with metadata
v0.10.0 - 2022-03-06
- Only the SP owning a Participant can delete it. That implies, that upon certificate change the simple deletion will not work. It is recommended to first index the participant, so that the new certificate is used, and than delete it with the new certificate.
- Added an Admin page to manually delete a participant without an owner check
- Showing metadata information on participant details, if the admin user is logged in
- Removed the "SMP implementations" page
- Added a possibility to hide or customize the "Contact us" page
- Changed the PD Client configuration properties, to start with
pdclient.and align the HTTP properties with SMP client configurationkeystore.typeis nowpdclient.keystore.typekeystore.pathis nowpdclient.keystore.pathkeystore.passwordis nowpdclient.keystore.passwordkeystore.key.aliasis nowpdclient.keystore.key.aliaskeystore.key.passwordis nowpdclient.keystore.key.passwordtruststore.typeis nowpdclient.truststore.typetruststore.pathis nowpdclient.truststore.pathtruststore.passwordis nowpdclient.truststore.passwordhttp.proxyHostis nowhttp.proxy.hosthttp.proxyPortis nowhttp.proxy.portproxy.usernameis nowhttp.proxy.usernameproxy.passwordis nowhttp.proxy.passwordconnect.timeout.msis nowhttp.connect.timeout.msrequest.timeout.msis nowhttp.request.timeout.mshttps.proxyHostis no longer supportedhttps.proxyPortis no longer supported
- Fixed the default search background image URL
v0.9.10 - 2022-02-24
- Prepare for internal cleanup to get rid of the legacy "deleted" flag
v0.9.9 - 2021-12-21
- Updated to Log4J 2.17.0 because of CVE-2021-45105 - see https://logging.apache.org/log4j/2.x/security.html
v0.9.8 - 2021-12-14
- Updated to Log4J 2.16.0 because of CVE-2021-45046 - see https://www.lunasec.io/docs/blog/log4j-zero-day/
v0.9.7 - 2021-12-10
- Updated to Log4J 2.15.0 because of CVE-2021-44228 - see https://www.lunasec.io/docs/blog/log4j-zero-day/
v0.9.6 - 2021-11-02
- Improved support for JSON API in Business Card
v0.9.5 - 2021-03-22
- Updated to ph-commons 10
- Updated to peppol-commons 8.4.0
- Improved web UI customizability
v0.9.4 - 2021-02-01
- Fixed initialization order issue
v0.9.3 - 2021-02-01
- Updated to ph-commons 9.5.4
- Updated to ph-dns 9.5.2
- Updated to Jersey 2.32
- Reduced lock contention
v0.9.2 - 2020-09-24
- Increased customizability
v0.9.1 - 2020-09-18
- Updated to Jakarta JAXB 2.3.3
v0.9.0 - 2020-09-16
- Updated to ph-commons 9.4.8
- Changed the way how the configuration system works
v0.8.8 - 2020-08-30
- Updated to ph-commons 9.4.7
- Updated to ph-oton 8.2.6
- Updated to peppol-commons 8.1.7
- Using Java 8 date and time classes for JAXB created classes
v0.8.7 - 2020-05-27
- Updated to ph-commons 9.4.4
- Updated to new Maven groupIds
- Improved logging
- Improved resilience on identifier handling for stored entries
v0.8.6 - 2020-02-19
- URL decoding participant identifiers on indexation
- Updated to ph-commons 9.4.0
v0.8.5 - 2020-02-16
- Finalized PEPPOL -> Peppol change
- Added registration date to the export data (see issue #45)
- Updated to peppol-commons 8.x
- Removed support for old PKI v2
- Made the identifier factory customizable to avoid duplicate entries
- Improved the internal Admin interface a bit
- Added possibility to automatically purge unwanted duplicate entries
- Updated the underlying UI libraries
- The lists of known document type IDs and process ID were updated
- Details about document types are now part of the export (see issue #46)
- Added the possibility to export search result as XML (see issue #43)
- Enforcing the
PDClientproxy configuration to be part ofPDHttpClientSettings - Improved internal error resilience
- Fixed a validation that broken the daily export because of invalid PD data
- Updated to ph-web 9.1.9
- Changed the internal
PDClientHTTP configuration API to useHttpClientSettings(backwards incompatible change) - The
PDClientnow checks for the key alias in a case insensitive manner (improved resilience)
v0.8.4 - 2020-01-24
- Updated to Jersey 2.30
- The Directory client has no more default truststore path and password
- The Directory client configuration can now be read from the path denoted by the environment variable
DIRECTORY_CLIENT_CONFIG - Updated the static texts changing
PEPPOLtoPeppol
v0.8.3 - 2020-01-08
- Added logo in the left top (using configuration property
webapp.applogo.image.path) - Setting
Content-LengthHTTP header for the downloads - Made FavIcons customizable
- Added rate limit for search API (using configuration property
rest.limit.requestspersecond)
v0.8.2 - 2019-10-14
- Added support to download all Business Cards as CSV
- Added support to download all Business Cards as XML but without the document types (see issue #42)
- Class
PDBusinessCardgot a default JSON representation - Updated to Jersey 2.29.1
- Added support to download all Participant IDs only as XML, JSON and CSV
v0.8.1 - 2019-07-29
- Updated to Jersey 2.29
PDClientConfigurationcan now be re-initialized during runtime- Known document type identifiers and process identifiers can now be used (see issue #13)
- Extended XML export to include the new document types (see issue #41)
- Added page to see the current Index Queue
v0.8.0 - 2019-06-27
- Renamed project
peppol-directorytophoss-directory - Maven artifact IDs changed from
peppol-directory*tophoss-directory* - Updated to
peppol-commons7.0.0 - Downgraded to Lucene 7.7.2
- Fixed an issue with
total-result-countand paging in the REST API (see issue #39) - Updated to Apache httpclient 4.5.9
- Updated to ph-oton 8.2.0
- Added a new internal page for importing identifiers
- The internal format for exporting participant IDs was updated to be used in the import
v0.7.2 - 2019-05-13
- Added new submodule
peppol-directory-searchapiwith basic elements for using the query API and the response documents - Updated default truststore of
peppol-directory-client - Updated to Lucene 8.1.0
v0.7.1 - 2019-03-17
- Added new method
PDBusinessCardHelper.parseBusinessCard - Updated to Lucene 8.0.0
v0.7.0 - 2018-12-02
- Added a link on the UI to download all business cards as XML
- Fixed the build timestamp property
- Fixed error when showing ReIndex entries of non-existing participants when using
ESensUrlProvider - Added the XML Schema for the API search results
- Added the XML Schema for the export data
- Added a page explaining the export data
- Requires ph-commons 9.2.0
- Updated UI to use Bootstrap 4.1
v0.6.2 - 2018-10-17
- If more hits are present than visible, it is displayed on the UI
- Made the available SML information objects customizable
- Removed the configuration item
sml.id- either fixed SMP or all configured SMLs are queried upon indexing - Updated to Apache Lucene 7.5
- Multilingual business entities are now supported via a new Business Card XML Schema - for Belgium
- The query API response document layout for XML was changed.
namehas now multiplicity 1..n instead of 1..1. - The query API response document layout for JSON was changed.
nameis now an array instead of astring. - Multiple parallel queries on the PD are possible.
v0.6.1 - 2018-06-04
- Avoid potential exception on invalid input parameters
- Updated to Jersey 2.27
- Updated to Apache Lucene 7.3
- Improved handling of multiple search parameters in name, geoinfo and additionalInfo
- Updated to peppol-commons 6.1.0
- Updated to ph-commons 9.1.0
- Introduced an internal "generic business card representation"
- An initial "export all business cards" was created
v0.6.0 - 2018-03-06
- Updated to ph-commons 9.0.1
- Updated to Apache Lucene 7.2.1
- Fixed some issues (as #30)
- Requires peppol-commons 6.0.1 for new OpenPEPPOL PKI v3
- Added support for trusting an arbitrary number of client certificate issuers (for the server only)
- Added support for configuring more than two truststores in pd.properties (for the server only)
- Added support for usage in the TOOP4EU project
- User interface texts can be changed from "PEPPOL Directory" to something else
- The PD client configuration now includes connection and request timeout, as well as proxy credentials
v0.5.1 - 2017-07-21
- Extended
PDClientto explicitly support a configurable truststore. A default truststore for the current setup is included. - PD client https hostname verification can now be
- PD client has now a custom exception callback to catch exceptions in the operations and handle them outside the client.
- Removed the JDK 6 PD client because the ECC certificates used are only supported by JDK 7 onwards. The old version is anyway in the Maven central repository.
v0.5.0 - 2017-07-12
- Updated release for
https://directory.peppol.euandhttps://test-directory.peppol.eu
My personal Coding Styleguide | It is appreciated if you star the GitHub project if you like it.