A Spring Boot 3.5.6 middleware service for secure batch file uploads with AWS S3 integration, designed for multi-tenant environments.
Data Forge Middleware provides a RESTful API for managing batch file uploads from client sites to AWS S3 storage. It features:
- Multi-tenant Architecture: Account-based isolation with site-level authentication
- Batch Upload Management: Track upload sessions with lifecycle states and metadata
- S3 Integration: Secure file storage with automatic retry and checksum validation
- Admin Portal: Auth0-secured (OAuth2 / JWT) endpoints for account and site management
- Observability: Structured JSON logging, metrics, and health checks
- PostgreSQL 16: Partitioned error logs and optimized queries
- Java 25 (Temurin/Corretto recommended)
- PostgreSQL 16+ with partitioning support
- AWS S3 or LocalStack for development
- Auth0 tenant (for the Admin/UI API and the frontend; see docs/local-auth0.md)
- Gradle 9.0+ (wrapper included)
This repo includes a DevContainer that starts infrastructure (PostgreSQL, Redis, LocalStack S3)
via docker-compose.dev.yml and provides a Java 25 + Node 20 dev environment.
- Configure backend env vars in
.env(see.env.example). - Configure frontend env vars in
frontend/.env.local(seefrontend/.env.local.example). - Open the folder
data-forge-middlewarein VS Code (not the parentbit-bifolder). - In VS Code: “Dev Containers: Reopen in Container”.
- Start backend:
./gradlew bootRun - Start frontend:
cd frontend && npm install && npm run dev
Notes:
- Auth is via Auth0. Keycloak has been fully removed (see Legacy: Keycloak).
- LocalStack bucket
dfm-uploadsis created automatically bydocker-compose.dev.yml. - Auth0 setup details:
docs/local-auth0.md - If backend fails with Flyway checksum mismatch, wipe local volumes:
docker-compose -f docker-compose.dev.yml down -vthenup -d.
Run infrastructure services in Docker and DFM from IDE for debugging:
# Start infrastructure (PostgreSQL, Redis, LocalStack)
./scripts/docker-dev.sh start
# Or manually
docker-compose -f docker-compose.dev.yml up -dThen in IntelliJ IDEA:
- Open Run Configuration
- Set Active profiles:
dev - Run or Debug the application
Infrastructure services:
- PostgreSQL: localhost:5432 (user:
postgres, password:postgres, database:dfm) - Redis: localhost:6379
- LocalStack S3: http://localhost:4566 (bucket: dfm-uploads)
Authentication is not part of the local stack — the backend validates Auth0-issued
JWTs against your tenant (AUTH0_* env vars, see docs/local-auth0.md).
Stop infrastructure:
./scripts/docker-dev.sh stopSee docker-compose.dev.yml for configuration details.
The easiest way to run the complete stack:
# Start all services (PostgreSQL, Redis, LocalStack S3, DFM Backend)
docker-compose up -d
# Check services are healthy
docker-compose ps
# View logs
docker-compose logs -f dfm-backendServices will be available at:
- Application API: http://localhost:8080
- Swagger UI: http://localhost:8080/swagger-ui.html
- PostgreSQL: localhost:5432 (database:
dfm) - LocalStack S3: http://localhost:4566
The backend requires AUTH0_* environment variables (see .env.example); roles come
from the https://api.dataforge.com/roles custom claim on the Auth0 access token —
there are no locally provisioned identity-provider users.
See docker/README.md for detailed Docker configuration.
git clone <repository-url>cd data-forge-middleware
./gradlew buildCreate PostgreSQL database:
CREATEDATABASEdataforge;
CREATEUSERdataforge_user WITH PASSWORD 'your-password';
GRANT ALL PRIVILEGES ON DATABASE dataforge TO dataforge_user;Copy example configuration:
cp src/main/resources/application-dev.yml.example src/main/resources/application-dev.ymlEdit application-dev.yml:
spring:
datasource:
url: jdbc:postgresql://localhost:5432/dataforgeusername: dataforge_userpassword: your-passwords3:
bucket:
name: dataforge-uploadsendpoint: http://localhost:4566 # LocalStackregion: us-east-1access-key: testsecret-key: testjwt:
secret: your-256-bit-secret-key-here-minimum-32-charsexpiration-minutes: 60batch:
timeout-minutes: 60docker run -d -p 4566:4566 -p 4571:4571 \
--name localstack \
-e SERVICES=s3 \
-e DEBUG=1 \
localstack/localstack:latest
# Create S3 bucket
aws --endpoint-url=http://localhost:4566 s3 mb s3://dataforge-uploads./gradlew flywayMigrate./gradlew bootRun --args='--spring.profiles.active=dev'Application starts on http://localhost:8080
Access interactive API documentation:
http://localhost:8080/swagger-ui.html
This is the only way a client obtains a token. There is no secret-based token endpoint —
the legacy POST /api/v1/auth/token was removed along with the rest of the v1 client surface.
curl -X POST http://localhost:8080/api/v1/device/authorize \
-H "Content-Type: application/json" \
-d '{ "siteName": "warehouse-01", "siteDescription": "Warehouse 01", "siteType": "DBF" }'Response:
{
"deviceCode": "GmRhmhcxhwAzkoEqiMEg_DnyEysNkuNhszIySk9eS",
"userCode": "WDJB-MJHT",
"verificationUri": "https://app.dataforge.com/device-verify",
"verificationUriComplete": "https://app.dataforge.com/device-verify?code=WDJB-MJHT",
"expiresIn": 900,
"interval": 5
}The account owner opens verificationUriComplete, authenticates via Auth0, and approves the
pending device. Auth0 authenticates the approver, not the device.
Poll no faster than interval seconds. Before approval the endpoint answers
authorization_pending.
curl -X POST http://localhost:8080/api/v1/device/token \
-H "Content-Type: application/json" \
-d '{"deviceCode": "GmRhmhcxhwAzkoEqiMEg_DnyEysNkuNhszIySk9eS"}'Response:
{
"siteId": "550e8400-e29b-41d4-a716-446655440000",
"siteName": "warehouse-01",
"accessToken": "eyJhbGciOiJIUzI1NiIs...",
"refreshToken": "kA9v...",
"accessTokenExpiresAt": "2026-01-01T13:00:00Z",
"refreshTokenExpiresAt": "2026-04-01T12:00:00Z",
"apiBaseUrl": "https://api.dataforge.com"
}Data does not go over REST. The client streams changes to the Delta v2 gRPC service
(default port 9090, contract in src/main/proto/delta-ingestion.proto), passing the access
token as a Bearer credential in gRPC metadata. See
docs/delta-client-v2-guide.md.
Refresh tokens rotate on every use: the presented token is revoked and a new one is returned. Presenting an already-rotated token is treated as a leak and revokes that rotation family.
curl -X POST http://localhost:8080/api/v1/device/auth/refresh \
-H "Content-Type: application/json" \
-d '{"refreshToken": "kA9v..."}'curl -X POST http://localhost:8080/api/v1/accounts \
-H "Authorization: Bearer <auth0-access-token>" \
-H "Content-Type: application/json" \
-d '{ "email": "user@example.com", "name": "Example User" }'curl -X POST http://localhost:8080/api/v1/accounts/{accountId}/sites \
-H "Authorization: Bearer <auth0-access-token>" \
-H "Content-Type: application/json" \
-d '{ "siteName": "warehouse-01", "displayName": "Warehouse 01" }'Keycloak was the original identity provider and has been fully replaced by Auth0.
The OAuth2 resource server is configured against the Auth0 issuer
(auth/config/Auth0SecurityConfig.java, shared/config/SecurityConfiguration.java,
spring.security.oauth2.resourceserver.jwt.issuer-uri = https://${auth0.domain}/),
and the frontend uses @auth0/auth0-react.
The Keycloak-era artefacts have been removed: the keycloak service and issuer in
docker-compose.prod.yml, the docker/keycloak/ realm files, the keycloak database and
role in docker/postgres/init.sql, and the dead src/main/resources/application-test.yml
(shadowed by src/test/resources/application-test.yml on the test classpath).
One deliberate exception remains in code: Auth0RoleConverter still falls back to
Keycloak's realm_access.roles claim, because the test harness itself
(config/TestSecurityConfig) mints tokens in that shape. Removing the fallback means
migrating the harness to Auth0-shaped claims first, and is tracked separately.
src/main/java/com/bitbi/dfm/
├── account/
│ ├── domain/ # Account aggregate
│ ├── application/ # AccountService, statistics
│ ├── infrastructure/ # JpaAccountRepository
│ └── presentation/ # AccountAdminController
├── site/
│ ├── domain/ # Site aggregate
│ ├── application/ # SiteService, event handlers
│ ├── infrastructure/ # JpaSiteRepository
│ └── presentation/ # SiteAdminController
├── batch/
│ ├── domain/ # Batch aggregate, BatchStatus
│ ├── application/ # BatchLifecycleService, timeout scheduler
│ ├── infrastructure/ # JpaBatchRepository
│ └── presentation/ # BatchController
├── upload/
│ ├── domain/ # UploadedFile, FileChecksum
│ ├── application/ # FileUploadService
│ ├── infrastructure/ # S3FileStorageService, config
│ └── presentation/ # FileUploadController
├── error/
│ ├── domain/ # ErrorLog (partitioned)
│ ├── application/ # ErrorLoggingService, export
│ ├── infrastructure/ # JpaErrorLogRepository, partition scheduler
│ └── presentation/ # ErrorLogController
├── auth/
│ ├── domain/ # JwtToken value object
│ ├── application/ # TokenService
│ ├── infrastructure/ # JwtTokenProvider, security config
│ └── presentation/ # AuthController
└── shared/
├── config/ # OpenAPI, Actuator, Metrics
├── exception/ # GlobalExceptionHandler, ErrorResponse
└── health/ # S3HealthIndicator
- accounts: User accounts with soft delete
- sites: Client sites with domain-based authentication
- batches: Upload sessions with lifecycle tracking
- uploaded_files: File metadata with S3 keys and checksums
- error_logs: Partitioned by month with JSONB metadata
Every zone-independent TIMESTAMP column in this schema holds a UTC wall clock, and nothing on
the way in or out converts it. A LocalDateTime in memory is therefore always that same UTC wall
clock — in an entity, after a raw-JDBC read, and in the DTOs, which serialize one with
toInstant(ZoneOffset.UTC).
- Producing — pass the zone explicitly:
LocalDateTime.now(ZoneOffset.UTC), andLocalDate.now(ZoneOffset.UTC)where a date picks a monthly partition oferror_logs. The zero-argumentnow()reads the JVM's default zone without saying so and is banned insrc/(TimestampProducerConventionTest). - Through Hibernate — an entity's
LocalDateTimefield is stored and read verbatim.spring.jpa.properties.hibernate.jdbc.time_zoneis deliberately not set: it would make Hibernate read a boundLocalDateTimeas wall clock in the JVM's zone and store a different instant, so the column would stop agreeing with every other path. Its absence is asserted, not merely left to habit. - Through raw JDBC (
JdbcTemplate/NamedParameterJdbcTemplate) — bind aLocalDateTimedirectly (JDBC 4.2) and read withrs.getObject(column, LocalDateTime.class). - From the database's own clock — some statements let PostgreSQL stamp the value, deliberately:
one clock across a fleet whose pods may disagree about the time (#245). PostgreSQL resolves
CURRENT_TIMESTAMPin the session's zone, and pgjdbc sets that zone from the JVM's default at connect time, so the statement must say UTC itself:CAST(current_timestamp AT TIME ZONE 'UTC' AS timestamp). JPQL has noAT TIME ZONE, so every such statement is native SQL — the queue markers inJpaChangelogSegmentRepository, the revocations inJpaRefreshTokenRepository, the activation touch inJpaAccountPluginRepository, and the Parquet Export catalog watermark. A bare clock function in production SQL fails the build (DatabaseClockConventionTest), and that the statements really do write UTC in a non-UTC session isDatabaseClockUtcIntegrationTest. The ban covers the test tree too (#287): a fixture that seeds a row from the session's clock is compared, off UTC, against values the application writes in UTC, so the test is red or green for a reason that is not its own. Two shapes are exempt there and each names why — a column that isTIMESTAMPTZ, where the bare form is the correct one and wrapping it would be the defect, and prose in a class with no database access that quotes the shape in order to talk about it. One population stays outside the ban and cannot reach production data: eleven applied migrations declare columnsDEFAULT CURRENT_TIMESTAMP— never edited, forward-only, and unreachable because every INSERT on those tables maps the column (new migrations are held to the convention).ENV TZ=UTCin theDockerfile(ContainerTimeZoneContractTest) is therefore belt-and-braces for the session zone rather than load-bearing; it still earns its place by making logs read in UTC.
A field held as Instant needs no rule: its binding stores the UTC wall clock in every zone and
never consulted hibernate.jdbc.time_zone at all.
Two guards hold this, and they are not interchangeable. TimestampProducerConventionTest is static
— it bans the zero-argument now() and asserts the setting's absence — and so gives the same answer
in every environment; it is what CI relies on.TimestampRoundTripIntegrationTest states the
property end to end (write through JPA, read the same column through raw JDBC, require equality, for
a LocalDateTime field and an Instant one) but runs in whatever zone the JVM is in: on a machine
outside UTC that makes it mutation-sensitive, while in CI, which runs in UTC, the conversion it
guards against would be the identity and the test would stay green. It deliberately does not
switch the JVM's default zone to change that, and #287 — which took the fixtures out of the
session-zone population — does not lift that restriction, only replaces its reason.
TimeZone.setDefault is a mutation of the whole process, shared with the background workers of
every cached context, and pgjdbc reads a connection's session zone at connect time, so a
connection opened during the window carries a non-UTC session back into the pool and outlives the
test that set it: the blast radius is the suite, not the method. What such a connection can still
reach is smaller than before but not empty — a TIMESTAMPTZ column read into a LocalDateTime
converts through that zone, and DatabaseClockUtcIntegrationTest measures the session clock on
purpose. Against that stands no gain this repository does not already have: SET LOCAL, undone by
the transaction it runs in, is how that test moves the zone safely. It buys nothing here, though
— the conversion this test guards is Hibernate's, keyed on the JVM's zone rather than the
session's — so the CI blind spot stands, and what closes it is the static guard.
java.sql.Timestamp is banned in src/ in either direction (Timestamp.valueOf, new Timestamp(…),
rs.getTimestamp(…)), because it is an instant and so silently carries the JVM's default zone.
rs.getTimestamp(col).toLocalDateTime() cancels that conversion for almost every value and is
therefore easy to mistake for a no-op, but not for a wall clock that falls in the JVM zone's DST
gap: that local time does not exist, Calendar resolves it forward, and the value comes back
shifted by the transition. RawTimestampReadConventionTest holds the ban; it carries no exemption
today (the one it had existed only to reapply the conversion #282 removed) and an entry added later
must name a shape, a count and a reason.
Until #282 the producing side was not uniform — parts of the code built a LocalDateTime with
LocalDateTime.now() (the JVM zone, which the Hibernate conversion then made correct) and parts
with LocalDateTime.now(ZoneOffset.UTC) (already UTC, which that same conversion applied a second
time). Removing the conversion is what let one producer be right on every path at once. On a UTC
JVM that conversion was the identity, so no stored row changed meaning and no data migration was
needed; rows written by a JVM outside UTC — local development only — shift by that JVM's offset.
- One Active Batch Per Site: Only one IN_PROGRESS batch allowed per site
- Concurrent Batch Limit: Maximum 5 active batches per account
- Batch Timeout: Batches auto-expire after 60 minutes (configurable)
- Cascade Deactivation: Deactivating account deactivates all sites
./gradlew test./gradlew integrationTest./gradlew test -PexcludeIntegrationThis runs the unit and contract suites without the Testcontainers integration suite. To run only a
specific contract class, use ./gradlew test --tests '*ContractTest'.
./gradlew jacocoTestReport
open build/reports/jacoco/test/html/index.htmlEvery Gradle Test task runs with maxHeapSize = "2g" (build.gradle.kts, tasks.withType<Test>).
Gradle's default is 512 MB, and the whole suite runs in one JVM — ~2470 tests, 444 classes and
24 cached Spring contexts held for the length of the run — so that default left no margin and CI
failed intermittently, naming a different, innocent test each time.
The value is measured rather than guessed. A -Xlog:gc pass over the full suite at a deliberately
generous 3 GiB never let G1 expand past 1014 MB, with the highest occupancy after a collection at
801 MB and the highest before one at 965 MB — so 1 GiB sits on the cliff and 512 MB was under it.
Re-measure before moving the value:
./gradlew test -PtestHeapLog # one GC log per Test task under build/reports/test-heap/One JVM is the intended shape. forkEvery would bound the accumulation by throwing away the
Spring TestContext cache — which is what makes 444 classes affordable — and it is the right tool
only for accumulation that has no ceiling; this accumulation is one context per distinct
configuration, a property of the test classes rather than of the test count. maxParallelForks
stays at 1 because the suite deliberately shares one PostgreSQL database across every context.
If the ceiling is reached anyway, -XX:+ExitOnOutOfMemoryError ends the JVM at the allocation that
failed instead of letting the error be caught and re-reported as some unrelated test's
BeanCreationException, and -XX:+HeapDumpOnOutOfMemoryError leaves the evidence in
build/reports/test-oom/. The build then fails as:
java.lang.OutOfMemoryError: Java heap space
Dumping heap to .../build/reports/test-oom/java_pid138960.hprof ...
Terminating due to java.lang.OutOfMemoryError: Java heap space
> Process 'Gradle Test Executor 4' finished with non-zero exit value 3
No test is named, because no test was at fault. The dump is deliberately not uploaded as a CI artifact — at this ceiling it can reach two gigabytes; reproduce locally instead.
Read the message before acting on it. The flag fires for every OutOfMemoryError the VM
raises, not only Java heap space: unable to create native thread and Metaspace end the worker
the same way and lose every remaining test result with it. Raising maxHeapSize is the remedy for
the first only — for native-thread exhaustion it makes matters worse. An OutOfMemoryError thrown
by ordinary code (a Mockito stub, as BatchParquetFinalizationIntegrationTest does) never reaches
this path and stays catchable.
curl http://localhost:8080/actuator/healthResponse:
{
"status": "UP",
"components": {
"db": { "status": "UP" },
"s3": { "status": "UP", "details": { "bucket": "dataforge-uploads" } },
"diskSpace": { "status": "UP" }
}
}curl http://localhost:8080/actuator/metricsCustom metrics:
batch.started- Total batches startedbatch.completed- Total batches completedbatch.failed- Total batches failedfiles.uploaded- Total files uploadederror.logged- Total errors logged
Structured JSON logging in production:
{
"@timestamp": "2024-01-01T12:00:00.000Z",
"level": "INFO",
"logger": "com.bitbi.dfm.batch.application.BatchLifecycleService",
"message": "Starting new batch",
"batchId": "550e8400-e29b-41d4-a716-446655440000",
"siteId": "123e4567-e89b-12d3-a456-426614174000",
"application": "data-forge-middleware"
}spring:
profiles:
active: proddatasource:
url: jdbc:postgresql://<rds-endpoint>:5432/dataforgehikari:
# Do not raise this without redoing the derivation beside the key in application.yml# (issue #161). The pool is per replica and the database's max_connections is not, so the# bound is (maxReplicas + maxSurge) x pool + an operator reserve <= max_connections, and at# the replica ceiling this deployment declares a server left at PostgreSQL's default 100# allows at most 12. Raising it needs `SHOW max_connections` on the actual server first —# BackgroundConnectionDemandTest fails until that number is recorded there.maximum-pool-size: 10minimum-idle: 5s3:
bucket:
name: prod-dataforge-uploadsregion: us-east-1# Uses IAM role credentials in productionlogging:
level:
root: INFOcom.bitbi.dfm: INFOFROM eclipse-temurin:21-jre-alpine
WORKDIR /app
COPY build/libs/*.jar app.jar
EXPOSE 8080
ENTRYPOINT ["java", "-jar", "app.jar"]Build and run:
./gradlew bootJar
docker build -t dataforge-middleware .
docker run -p 8080:8080 \
-e SPRING_PROFILES_ACTIVE=prod \
-e SPRING_DATASOURCE_URL=jdbc:postgresql://db:5432/dataforge \
dataforge-middleware- Follow Java 25 conventions
- Use Lombok for boilerplate reduction
- Domain-driven design principles
- Package by layered feature (PbLF)
- Create feature branch:
git checkout -b feature/my-feature - Write tests for new functionality
- Ensure all tests pass:
./gradlew test - Update documentation as needed
- Submit PR with clear description
Proprietary - Bit BI
For issues or questions:
- Email: support@bitbi.com
- Documentation: https://docs.dataforge.bitbi.com