Commits · a5202bc968cf3ca5b64c623b9271f76e8fa02211 · libremedia / Tethys / FFmpeg

Nov 09, 2015

swresample/resample: improve bessel function accuracy and speed · a5202bc9

This improves accuracy for the bessel function at large arguments, and this in turn
should improve the quality of the Kaiser window. It also improves the
performance of the bessel function and hence build_filter by ~ 20%.
Details are given below.

Algorithm: taken from the Boost project, who have done a detailed
investigation of the accuracy of their method, as compared with e.g the
GNU Scientific Library (GSL):
http://www.boost.org/doc/libs/1_52_0/libs/math/doc/sf_and_dist/html/math_toolkit/special/bessel/mbessel.html.
Boost source code (also cited and licensed in the code):
https://searchcode.com/codesearch/view/14918379/

.

Accuracy: sample values may be obtained as follows. i0 denotes the old bessel code,
i0_boost the approach here, and i0_real an arbitrary precision result (truncated) from Wolfram Alpha:
type "bessel i0(6.0)" to reproduce. These are evaluation points that occur for
the default kaiser_beta = 9.

Some illustrations:
bessel(8.0)
i0      (8.000000) = 427.564115721804739678191254
i0_boost(8.000000) = 427.564115721804796521610115
i0_real (8.000000) = 427.564115721804785177396791

bessel(6.0)
i0      (6.000000) = 67.234406976477956163762428
i0_boost(6.000000) = 67.234406976477970374617144
i0_real (6.000000) = 67.234406976477975326188025

Reason for accuracy: Main accuracy benefits come at larger bessel arguments, where the
Taylor-Maclaurin method is not that good: 23+ iterations
(at large arguments, since the series is about 0) can cause
significant floating point error accumulation.

Benchmarks: Obtained on x86-64, Haswell, GNU/Linux via a loop calling
build_filter 1000 times:
test: fate-swr-resample-dblp-44100-2626

new:
995894468 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1029719302 decicycles in build_filter(loop 1000),     512 runs,      0 skips
984101131 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

old:
1250020763 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1246353282 decicycles in build_filter(loop 1000),     512 runs,      0 skips
1220017565 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

A further ~ 5% may be squeezed by enabling -ftree-vectorize. However,
this is a separate issue from this patch.

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

a5202bc9

swresample: allow double precision beta value for the Kaiser window · 1bed09a3

Ganesh Ajjanagadde authored 9 years ago


Kaiser windows inherently don't require beta to be an integer. This was
an arbitrary restriction. Moreover, soxr does not require it, and in
fact often estimates beta to a non-integral value.

Thus, this patch allows greater flexibility for swresample clients.
Micro version is updated.

Reviewed-by: Derek Buitenhuis <derek.buitenhuis@gmail.com>
Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

1bed09a3

Nov 06, 2015

swresample/resample: speed up build_filter for Blackman-Nuttall filter · c8780822

Ganesh Ajjanagadde authored 9 years ago


This uses the trigonometric double and triple angle formulae to avoid
repeated (expensive) evaluation of libc's cos().

Sample benchmark (x86-64, Haswell, GNU/Linux)
test: fate-swr-resample-dblp-44100-2626
old:
1104466600 decicycles in build_filter(loop 1000),     256 runs,      0 skips
1096765286 decicycles in build_filter(loop 1000),     512 runs,      0 skips
1070479590 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

new:
588861423 decicycles in build_filter(loop 1000),     256 runs,      0 skips
591262754 decicycles in build_filter(loop 1000),     512 runs,      0 skips
577355145 decicycles in build_filter(loop 1000),    1024 runs,      0 skips

This results in small differences with the old expression:
difference (worst case on [0, 2*M_PI]), argmax 0.008:
max diff (relative): 0.000000000000157289807188
blackman_old(0.008): 0.000363951585488813192382
blackman_new(0.008): 0.000363951585488755946507

These are judged to be insignificant for the performance gain. PSNR to
reference file is unchanged up to second decimal point for instance.

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

c8780822

Nov 04, 2015

swresample/resample: speed up build_filter by 50% · 9bec6d71

Ganesh Ajjanagadde authored 9 years ago

This speeds up build_filter by ~ 50%. This gain should be pretty
consistent across all architectures and platforms.

Essentially, this relies on a observation that the filters have some
even/odd symmetry that may be exploited during the construction of the
polyphase filter bank. In particular, phases (scaled to [0, 1]) in [0.5, 1] are
easily derived from [0, 0.5] and expensive reevaluation of function
points are unnecessary. This requires some rather annoying even/odd
bookkeeping as can be seen from the patch.

I vaguely recall from signal processing theory more general symmetries allowing even greater
optimization of the construction. At a high level, "even functions"
correspond to 2, and one can imagine variations. Nevertheless, for the sake
of some generality and because of existing filters, this is all that is
being exploited.

Currently, this patch relies on phase_count being even or (trivially) 1,
though this is not an inherent limitation to the approach. This
assumption is safe as phase_count is 1 << phase_bits, and is hence a
power of two. There is no way for user API to set it to a nontrivial odd
number. This assumption has been placed as an assert in the code.

To repeat, this assumes even symmetry of the filters, which is the most common
way to get generalized linear phase anyway and is true of all currently
supported filters.

As a side note, accuracy should be identical or perhaps slightly better
due to this "forcing" filter symmetries leading to a better phase
characteristic. As before, I can't test this claim easily, though it may
be of interest.

Patch tested with FATE.

Sample benchmark (x86-64, Haswell, GNU/Linux):

test: swr-resample-dblp-44100-2626

new:
527376779 decicycles in build_filter(loop 1000), 256 runs, 0 skips
524361765 decicycles in build_filter(loop 1000), 512 runs, 0 skips
516552574 decicycles in build_filter(loop 1000), 1024 runs, 0 skips

old:
974178658 decicycles in build_filter(loop 1000), 256 runs, 0 skips
972794408 decicycles in build_filter(loop 1000), 512 runs, 0 skips
954350046 decicycles in build_filter(loop 1000), 1024 runs, 0 skips

Note that lower level optimizations are entirely possible, I focussed on
getting the high level semantics correct. In any case, this should
provide a good foundation.

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

9bec6d71

Oct 28, 2015
- swr: do not reject channel layouts that use channel 63 · 80580bb2
  wm4 authored 9 years ago
  
  Channel layouts are essentially uint64_t, and every value is valid.
  80580bb2
Oct 25, 2015

all: add const-correctness to qsort comparators · c7131762

Ganesh Ajjanagadde authored 9 years ago


This adds const-correctness when needed for the comparators.

Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

c7131762

Oct 22, 2015

avfilter,swresample,swscale: use fabs, fabsf instead of FFABS · 8507b98c

Ganesh Ajjanagadde authored 9 years ago

It is well known that fabs and fabsf are at least as fast and sometimes
faster than the FFABS macro, at least on the gcc+glibc combination.
For instance, see the reference:
http://patchwork.sourceware.org/patch/6735/

.
This was a patch to glibc in order to remove their usages of a macro.

The reason essentially boils down to fabs using the __builtin_fabs of
the compiler, while FFABS needs to infer to not use a branch and to
simply change the sign bit. Usually the inference works, but sometimes
it does not. This may be easily checked by looking at the asm.

This also has the added benefit of reducing macro usage, which has
problems with side-effects.

Note that avcodec is not handled here, as it is huge and
most things there are integer arithmetic anyway.

Tested with FATE.

Reviewed-by: Clément Bœsch <u@pkh.me>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

8507b98c

Oct 16, 2015

swresample/swresample_internal: add av_warn_unused_result · ef62f573

Ganesh Ajjanagadde authored 9 years ago


This will trigger a few warnings that need to be fixed.

Reviewed-by: Michael Niedermayer <michael@niedermayer.cc>
Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>

ef62f573

Oct 15, 2015
- swresample: slightly nicer debug output for auto matrix · cdf4a13f
  wm4 authored 9 years ago
  
  This is the matrix that will be used for up/downmixing.
  cdf4a13f
Oct 10, 2015

doc/resampler, swresample/options: use proper capitalization · f3fc103c

Ganesh Ajjanagadde authored 9 years ago


Proper names should be capitalized in all user facing API as far as
possible. The option names themselves have not been changed since:
1. We consistently keep option names in lower case.
2. Changing them would break existing scripts.
3. I suspect that we want to be similar to Sox and its relevant options.

The converse is also true: improper names should not be capitalized
generally.

Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

f3fc103c

Oct 07, 2015
- swresample/resample: manually unroll the main loop in bessel() · 1bc873ac
  Michael Niedermayer authored 9 years ago
  
  About 10% faster Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
  1bc873ac
- swresample/resample: merge first iteration into init in bessel() · 6024c865
  Michael Niedermayer authored 9 years ago
  
  speedup of about 1% Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
  6024c865
Oct 02, 2015

x86/audio_convert: fix clobbering of xmm registers · acdd6725

James Almer authored 9 years ago


Reviewed-by: Michael Niedermayer <michaelni@gmx.at>
Signed-off-by: James Almer <jamrial@gmail.com>

acdd6725

Sep 27, 2015
- swresample/dither_template: Add missing license header · 7d636d02
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
  7d636d02
Sep 03, 2015
- swresample/swresample: Fix integer overflow in seed calculation · 32f53958
  Michael Niedermayer authored 9 years ago
  
  Fixes CID1322333 Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
  32f53958
Aug 30, 2015
- swresample/swresample-test: Make layouts static const · fb42e775
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>
  fb42e775
Aug 23, 2015

swresample/dither: use integer arithmetic · 24e6729a

Ganesh Ajjanagadde authored 9 years ago

This fixes a -Wabsolute-value reported by clang 3.5+ complaining about misuse of fabs() for integer absolute value.
An additional benefit is the removal of floating point calculations.

Signed-off-by: Ganesh Ajjanagadde <gajjanagadde@gmail.com>
Signed-off-by: Michael Niedermayer <michael@niedermayer.cc>

24e6729a

Aug 03, 2015

x86: move XOP emulation code back to x86inc · 5750d6c5

James Almer authored 9 years ago


Only two functions that use xop multiply-accumulate instructions where the
first operand is the same as the fourth actually took advantage of the macros.

This further reduces differences with x264's x86inc.

Reviewed-by: Ronald S. Bultje <rsbultje@gmail.com>
Signed-off-by: James Almer <jamrial@gmail.com>

5750d6c5

Jul 26, 2015
- swresample/x86: add missing colon to labels · f37a5dcb
  James Almer authored 9 years ago
  
  Silences warnings with Nasm Signed-off-by: James Almer <jamrial@gmail.com>
  f37a5dcb
Jul 16, 2015
- lswr: Allow 64 channels internally. · a77401e1
  Carl Eugen Hoyos authored 9 years ago
  
  a77401e1
Jun 22, 2015
- swr: Remember previously set int_sample_format from user · d4325b2f
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  d4325b2f
- swresample/swresample: Clear delayed_samples_fixup in clear_context() · 0dd2790d
  Michael Niedermayer authored 9 years ago
  
  This probably makes no difference but its more proper Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  0dd2790d
Jun 21, 2015

swresample: soxr implementation for swr_get_out_samples() · c70c6be2
Rob Sykes authored 9 years ago
```
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
```
c70c6be2
swresample/swresample: Print used int_sample_fmt · 5de3a589
Michael Niedermayer authored 9 years ago
```
Suggested-by: wm4
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
```
5de3a589

swresample: Choose 16bit internally only if input and output is 16bit or less · 49776924

Michael Niedermayer authored 9 years ago


or if no rematrix and no resampling is performed and the input is 16bit
note reampling and rematrix itself always use more than 16bit internally
the "internal" sampling format is the format between these steps

Its unlikely the difference from this commit is audible in any case
unless there is some bug either before or after the change.
but multiple people prefer this and it slightly improves the precission
of computations.

Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

49776924

Jun 08, 2015
- swr: Fix ASSERT_LEVEL warning · 56f0fe6b
  Michael Niedermayer authored 9 years ago
  
  Found-by: cehoyos Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  56f0fe6b
Jun 06, 2015
- swresample: fix initilaize/initialize typo · c5a08956
  Clément Bœsch authored 9 years ago
  
  c5a08956
Jun 04, 2015

swresample/resample: fix typos · b1436148

Michael Niedermayer authored 9 years ago


Found-by: wm4 <nfxjfg@googlemail.com>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

b1436148

swresample/swresample: Cleanup on init failure. · c3f87f75

Michael Niedermayer authored 9 years ago


This avoids leaks if the user doest call swr_close() after a failed init

Found-by: James Almer <jamrial@gmail.com>
Reviewed-by: James Almer <jamrial@gmail.com>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

c3f87f75

swresample: Add swr_get_out_samples() · cc17b43d

Michael Niedermayer authored 9 years ago


Previous version reviewed-by: Pavel Koshevoy <pkoshevoy@gmail.com>
Previous version reviewed-by: wm4 <nfxjfg@googlemail.com>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

cc17b43d

libswresample/rematrix: Check for malloc errors · 52acd22a
Michael Niedermayer authored 9 years ago
```
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
```
52acd22a

Jun 03, 2015

swresample/dither: check memory allocation · 196b885a

Ganesh Ajjanagadde authored 9 years ago


check memory allocation in swri_get_dither()

Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

196b885a

Jun 02, 2015
- swresample: Check the return value of resampler->init() · 02915602
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  02915602
May 31, 2015

x86: check for AV_CPU_FLAG_AVXSLOW where useful · c16e99e3

James Almer authored 9 years ago


Signed-off-by: James Almer <jamrial@gmail.com>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

c16e99e3

May 13, 2015

swr: fix alignment issue caused by 8ch sse functions · adb7372f

Rainer Hochecker authored 9 years ago

Fix crash when doing 8 ch conversion from apps compiled with MSVS
Thanks to Ronald for giving this hint:
https://ffmpeg.org/pipermail/ffmpeg-devel/2015-May/173049.html



Reviewed-by: "Ronald S. Bultje" <rsbultje@gmail.com>
Signed-off-by: Michael Niedermayer <michaelni@gmx.at>

adb7372f

May 06, 2015
- swresample/dither_template: Do not define macro functions to nothing · 223a8598
  Michael Niedermayer authored 9 years ago
  
  This avoids potential warnings Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  223a8598
Apr 12, 2015
- swresample/swresample-test: Randomly wipe out channel counts · ff50b1b1
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  ff50b1b1
- swresample: Check channel layouts and channels against each other and print... · 3c77bb5f
  Michael Niedermayer authored 9 years ago
  
  swresample: Check channel layouts and channels against each other and print human readable error messages Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  3c77bb5f
- swresample: Allow reinitialization without ever setting channel layouts · 80a28c75
  Michael Niedermayer authored 9 years ago
  
  80a28c75
- swresample: Allow reinitialization without ever setting channel counts · d7b9cb2f
  Michael Niedermayer authored 9 years ago
  
  Signed-off-by: Michael Niedermayer <michaelni@gmx.at>
  d7b9cb2f