3 ms·
The perf part of the tests just seems to be a microbenchmark for seeing how fast the various frameworks can parse a 30000x300 dict of strings representing numbe
by no_circuit 4y ago
The perf part of the tests just seems to be a microbenchmark for seeing how fast the various frameworks can parse a 30000x300 dict of strings representing numbers [1].
If that is all one's application does, and can use your library in their organization/team, that's great. However a 2-3x performance boost for the parsing stage for a use case like an API call might not matter when that could be overshadowed by validation and/or upstream API calls. A realistic app would likely use a validation library like Pydantic's [2] to throw a custom typed-error that can be processed, e.g., localization, before returning it downstream.
[1] https://github.com/ltworf/typedload/blob/37c72837e0a8fd5f3502182e8d7716f13045e4d7/perftest/load%20big%20dictionary.py#L70 https://github.com/ltworf/typedload/blob/37c72837e0a8fd5f350...
[2] https://docs.pydantic.dev/usage/validators/ https://docs.pydantic.dev/usage/validators/
- LtWorf 4y ago> a microbenchmark for seeing how fast the various frameworks can parse a 30000x300 dict of strings representing numbers [1]. Did you somehow miss all the other tests, even thought they are higher on the page? The important test is shown first: loading objects. I'm not trying to benchmark my wifi or my disk. The IO time is not included there on purpose. Of course a bigger application that does other things wouldn't see this huge difference in performance. But I'm testing the performance of a library here. typedload can use validators, but since it doesn't reimplement attrs/dataclass, I had no interest in testing those… it'd be a race between libraries I didn't write. typedload's exceptions contain enough detail to tell the end user what went wrong. when encountering errors in lists, typedload slows down due to keeping compatibility with python 3.7. However in my test with loading objects, despite the slowdown it remains faster than pydantic.
- no_circuit 4y agoI read all the perftests in the repo. I think they nearly all parse a structure that contains a repetition of the same or similar thing a couple hundred thousand times times and the timing function returns the min and max of 5 attempts. I just picked one example for posting. Not a Python expert, but could the Pydantic tests be possibly not realistic and/or misleading because they are using kwargs in __init__ [1] to parse the object instead of calling the parse_obj class method [2]? According to some PEPs [3], isn't Python creating a new dictionary for that parameter which would be included in the timing? That would be unfortunate if that accounted for the difference. Something else I think about is if a performance test doesn't produce a side effect that is checked, a smart compiler or runtime could optimize the whole benchmark away. Or too easy for the CPU to do branch prediction, etc. I think I recall that happening to me in Java in the past, but probably not happened here in Python. [1] https://github.com/ltworf/typedload/blob/37c72837e0a8fd5f3502182e8d7716f13045e4d7/perftest/load%20big%20dictionary.py#L39 https://github.com/ltworf/typedload/blob/37c72837e0a8fd5f350... [2] https://docs.pydantic.dev/usage/models/#helper-functions https://docs.pydantic.dev/usage/models/#helper-functions [3] https://peps.python.org/pep-0692/ https://peps.python.org/pep-0692/
- LtWorf 4y agoHow would you implement a benchmark then? The kwargs thing is true… but I didn't design the API of pydantic. And it only happens on the small top level dictionary, not on all of them (unless it internally does it all the time). Java has JIT, I agree that in that case keeping the output value is important. cpython isn't that smart. I honestly never tried to benchmark using pypy. I guess it could be interesting to try that.