콘텐츠로 건너뛰기
블로그로 돌아가기

영어 원본

Engineering notes · Novel-view synthesis

Evaluating Indoor 3D Reconstruction: Dr Johnson and Playroom

MakeWorlds reconstructed Dr Johnson and Playroom from photographs using its own camera estimates. In the recorded Mac runs, both scenes return higher PSNR and SSIM and lower LPIPS than the corresponding values reported in the original 3DGS paper. The image pairs show what the reconstructions retain in practice: cabinet profiles, repeating rug patterns, and the layered contents of a bookshelf.

MakeWorlds team · Tested

1. Results across the full test set

The averages below cover all 16 test views in Dr Johnson and all 14 in Playroom. In both scenes, all three metrics are favorable relative to the reported 3DGS values. Four image pairs later in the article show the visible detail behind those results.

The tables use the paper-comparison protocol from our published evaluation, without substituting color-corrected scores. Reference values are the 30,000-step results in the original 3DGS appendix. Camera estimates and test-view sets differ, so the deltas provide context against the literature rather than a controlled comparison. [1] Comparison conditions and scope

Dr Johnson · Mac · Full test-set means · Paper-comparison protocol
Δ = MakeWorlds − paper. Higher PSNR / SSIM and lower LPIPS are better.
ResultPSNR ↑ / dBSSIM ↑LPIPS ↓
3DGS paper (reported)28.7660.89900.2440
MakeWorlds · Mac29.180Δ +0.4140.9124Δ +0.01340.2215Δ −0.0225
Playroom · Mac · Full test-set means · Paper-comparison protocol
Δ = MakeWorlds − paper. Higher PSNR / SSIM and lower LPIPS are better.
ResultPSNR ↑ / dBSSIM ↑LPIPS ↓
3DGS paper (reported)30.0440.90600.2410
MakeWorlds · Mac30.385Δ +0.3410.9148Δ +0.00880.2264Δ −0.0146

The direction is consistent across both scenes: higher PSNR and SSIM, lower LPIPS. These measures describe pixel error, local image structure, and perceptual distance. The comparisons below connect that scene-level assessment to specific features, including cabinet edges, floorboard seams, and repeated patterns.

2. From photographs to evaluated views

Dr Johnson and Playroom are public Deep Blending scenes also used in the original 3DGS evaluation. All MakeWorlds images and scores here come from the Mac runs in the published October 1, 2026 batch. [2]

Dataset and viewsDr JohnsonPlayroom
Input photos263225
Test views1614
Views shown here22
Image size (pixels)1332 × 8751264 × 832

We ran the Medium preset at source resolution for 30,000 optimization steps on an Apple M4 Max. MakeWorlds estimated cameras from the input images, trained the scene, and rendered the test views. Those views were fixed before training; their images were excluded from scene optimization and used as references for evaluation.

The four illustrated views cover cabinetry, the play area, rug patterns, and a crowded bookshelf. We selected them for those features, not for the highest view scores. Each reference/render pair uses matching display orientation and cropping, making it possible to check individual features against the scene-level results.

3. Dr Johnson: from the cabinet outline to its internal framing

The reconstruction retains detail at several scales within the same view. The dark cabinet keeps its overall profile as well as the pointed arches and horizontal divisions of its doors. The fireplace mantel, side supports, and wall moldings provide further contours to follow. Together, these features show how the reconstructed furnishings relate to the surrounding room.

MakeWorlds render: The wooden cabinet, white fireplace, and green walls in Dr Johnson
Original photograph: The wooden cabinet, white fireplace, and green walls in Dr Johnson
Original photoMakeWorlds render

Figure 1. Dr Johnson, Mac, view-000040. The cabinet’s arches and internal divisions remain distinct alongside its outer profile and the fireplace moldings. Reference photo on the left; MakeWorlds render on the right.

The useful evidence here is the combination of a recognizable room layout and the smaller structures inside it. Dragging the divider across the cabinet brings those correspondences into view. The bright window area and some fine edges still differ visibly from the reference.

4. Playroom: color regions, separate objects, and floor texture

Playroom offers several scales of detail in one image. The reconstruction reproduces the broad rug stripes and foam-mat lettering, while the scattered toys keep identifiable shapes and positions. Away from the colorful play area, floorboard seams and knots provide another set of correspondences. The image retains information across broad color regions, object boundaries, and surface texture.

MakeWorlds render: The striped rug, toys, and wooden floor in Playroom
Original photograph: The striped rug, toys, and wooden floor in Playroom
Original photoMakeWorlds render

Figure 2. Playroom, Mac, view-000184. Rug stripes, foam-mat lettering, toy outlines, and floorboard seams provide checks at progressively finer scales.

5. Closer inspection: repeated patterns and bookshelf contents

On the Dr Johnson rug, the border motifs remain arranged in a continuous sequence, and the shapes and positions of the larger central patterns can be matched to the photograph. Playroom preserves the widths and colors of neighboring book spines, the angles of the binders, and the layers of stacked objects. On the upper shelf, the green book’s title remains legible. Both pairs below use the same crop and magnification for reference and render, so these features can be compared directly.

MakeWorlds render: The rug, chairs, and wooden floor in Dr Johnson
Original photograph: The rug, chairs, and wooden floor in Dr Johnson
Original photoMakeWorlds render

Figure 3a. Dr Johnson, view-000248. The rug’s border and central motifs remain individually identifiable. Some finer pattern elements and chair-back edges are softer.

MakeWorlds render: Bookshelves, binders, and cupboard doors in Playroom
Original photograph: Bookshelves, binders, and cupboard doors in Playroom
Original photoMakeWorlds render

Figure 3b. Playroom, view-000089. Book spines and binders retain their separate outlines, colors, and arrangement. Larger title lettering remains readable; smaller text shows varying degrees of softening.

6. Connecting visible detail to the metrics

The examples move from room layout and broad color regions to object boundaries, repeated motifs, and individual letter strokes. Quantitative evaluation summarizes image differences in terms of pixel error, local structure, and perceptual features. Using both gives the reader numerical results to compare and concrete image features to inspect.

PSNR ↑
PSNR expresses pixel mean-squared error on a logarithmic scale. Higher values indicate lower error under the chosen comparison conditions. Image alignment, brightness, and color differences all affect it.
SSIM ↑
SSIM compares local luminance, contrast, and structure. Higher values indicate a closer match in those image properties, adding a structural view of similarity alongside pixel error. [3]
LPIPS ↓
LPIPS measures distance between deep image features; lower values indicate a closer match to the reference. The feature network and preprocessing affect the values, so comparisons across implementations require a consistent protocol. [4]

The public 3D Gaussian Splatting model provides a useful way to think about these images. A scene is represented by Gaussian primitives, which are projected into the image and blended in depth order. Omitting the background term gives: [1]

C(u, d) = ∑i Ti(u) · αi(u) · ci(d)Ti(u) = ∏j<i [1 − αj(u)]

Here, u is the pixel position and d is the viewing direction. Alpha combines the projected Gaussian footprint with opacity; T is the transmittance remaining after the preceding Gaussians. Several contributions form each pixel, making coverage, occlusion, and local color variation relevant to the final image. Cabinet framing, toy edges, and rug patterns provide concrete places to inspect those effects.

7. Conclusions and scope

Across these two indoor cases, MakeWorlds preserves room layout, furniture structure, and fine surface detail, alongside favorable results across the full test sets. The evidence is visible at several scales: a room’s overall arrangement, a cabinet’s internal framing, and patterns or lettering within the scene. These runs demonstrate detailed indoor reconstruction from a workflow that starts with photographs and estimates its own cameras.

These two indoor cases were selected from the published evaluation because all three metrics were favorable relative to the paper references. MakeWorlds uses its own camera estimates, and the test-view sets and complete pipelines were not matched to the paper’s experiment. The differences offer context against reported literature values; they do not establish a controlled ranking or isolate the contribution of an individual component.

The conclusions concern image fidelity at the tested viewpoints in these scenes and settings. Measurement applications would also need independent geometric ground truth, dimensional error tests, and surface checks. Walkthroughs would need inspection during camera motion for flicker, occlusion transitions, and newly exposed regions. Those tests are outside this article.

We report one recorded evaluation batch without run-to-run variance or confidence intervals. The deltas are observations, with no claim of statistical significance.

Test conditions and sources
Input
263 input photos for Dr Johnson; 225 for Playroom.
Settings
Medium preset, source resolution, 30,000 steps. Images and scores are from the Mac runs in the same evaluation batch.
Images
Dr Johnson: view-000040 and view-000248. Playroom: view-000184 and view-000089. Views were chosen to inspect layout, edges, and texture, rather than to showcase the highest scores. Each render is paired with its reference photo; display orientation and crop are matched.

Dataset: Deep Blending. Deep Blending

References

The rendering model and metric definitions follow the sources below. The experimental results are drawn from MakeWorlds’ published scene evaluations.

  1. Kerbl et al. (2023). 3D Gaussian Splatting for Real-Time Radiance Field Rendering.
  2. Hedman et al. (2018). Deep Blending for Free-Viewpoint Image-Based Rendering.
  3. Wang et al. (2004). Image Quality Assessment: From Error Visibility to Structural Similarity.
  4. Zhang et al. (2018). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric.

Explore the three-platform benchmark results · Plan a photo capture