NeztoLästid: 4 min
Master Thesis Project - 2027
Nezto is offering a master's thesis on using image data to improve our AI-driven property valuation engine, the Nezto Vision Engine™, exploring computer vision techniques to capture signal tabular data misses, fused into a live production model. GPU compute, real listing data, and guidance from our ML team and Modulai.

Stockholm/Gothenburg
Background
Nezto is building the Nezto Vision Engine™, a state-of-the-art automated valuation model (AVM) that combines tabular data, computer vision and explainable machine learning to value residential properties at scale. We're offering a master's thesis project on using image data to further improve the accuracy of our valuation engine, with guidance from our own ML team and the ML engineers at Modulai.
Automated valuation models traditionally rely on tabular data: living area, number of rooms, location, construction year and historical transactions. Two apartments with near-identical records can still differ substantially in market value because of condition, renovation standard, light, layout and view. Much of that residual signal is present in listing photographs, floor plans and aerial imagery, but is rarely exploited beyond coarse heuristics. Early work showed that a learned "luxury level" derived from interior and exterior photos, combined with metadata, can outperform established metadata-only estimates (Poursaeed et al., 2017), and later studies confirm that visual features add predictive power on top of strong tabular baselines (Kostic & Jevremović, 2021).
This thesis treats the image side as a representation and integration problem: which visual representations carry the signal that tabular features miss, and how should they be fused into a production valuation model without hurting robustness, calibration or explainability? The candidate representations span the full toolbox — image classification (room type, condition, renovation standard), semantic segmentation (materials, surfaces, greenery, floor-plan geometry), object detection (fireplaces, appliances, balconies) and general-purpose embeddings from pretrained vision or vision-language backbones.
Core idea: Explore image model approaches to represent image information as efficiently as possible while keeping the model explainable, and investigate how such methods can improve valuation accuracy for apartments and houses. The methodological contribution is the comparison and fusion strategy; the applied contribution is a validated improvement in a system that is actually shipped and already used in production by broker chains, property developers and property platforms.
Example Directions
- Systematic comparison of representation families: classification heads, segmentation masks, detected objects and raw embeddings on equal footing, measured as marginal accuracy over our strong tabular baseline
- Off-the-shelf vision-language models used as attribute extractors versus purpose-trained models and learned embeddings: accuracy, cost and latency per valuation
- Fusion architecture: late fusion of pooled embeddings into a gradient-boosted model versus end-to-end multimodal training, including how to aggregate a variable number of images per property
- Weak supervision from price residuals: learning image representations directly against the part of the price our tabular model cannot explain, instead of relying on generic pretrained features
- Floor plans as structured input: extracting layout, room adjacency and geometry, and testing whether structure beats appearance
- Robustness to presentation: staged photos, wide-angle lenses, HDR and photographer differences separating property quality from marketing quality
- Uncertainty and explainability: does image data mainly shift the point estimate or tighten prediction intervals, and can per-image contributions be attributed in a way a human valuer would accept?
ML Techniques and Tools
- Python, PyTorch, Git, Hugging Face
- CNN and vision-transformer backbones; pretrained embeddings (e.g. CLIP, DINOv2) and vision-language models
- Semantic segmentation and object detection (e.g. SAM, Mask R-CNN, YOLO-family models)
- Multimodal fusion and gradient boosting (LightGBM/XGBoost) over combined tabular and image features
- Explainability (SHAP, attention and saliency maps) and uncertainty quantification (quantile regression, conformal prediction)
- Cloud GPU compute, experiment tracking and evaluation on real data
Why Nezto
You'll work with large-scale real listing data, GPU compute, and close guidance from both our ML team and Modulai's ML engineers, on a model that is actually shipped and used in production. It's also a domain that's unusually easy to relate to: everyone lives somewhere.
How to Apply
Send your application to hello@nezto.com.
References
- Poursaeed et al., Vision-based Real Estate Price Estimation, 2017. arXiv:1707.05489 — https://arxiv.org/abs/1707.05489
- Kostic & Jevremović, What Image Features Boost Housing Market Predictions?, 2021. arXiv:2107.07148 — https://arxiv.org/abs/2107.07148
- Zillow neural network estimate — https://www.zillow.com/news/building-the-neural-zestimate


